The archive, mapped
Every blog post and case study on the site with the principles and playbooks each piece argues for. Filter by any principle or playbook to see just the evidence behind it.
Matching the filter
7 itemsHow I evolved a small AMD GPU cluster into a CRD-driven inference platform with an OpenAI-compatible boundary, evidence-based model promotion, and safe rollouts.
- SLOs for Inference: Latency, Errors, SaturationDec 29, 2025 · 5 min
How I define request, cold-start, queue, and model-quality objectives for an inference service without confusing causes with user outcomes.
- Standing Up a GPU-Ready Private AI Platform (Harvester + K3s + Flux + GitLab)Dec 29, 2025 · 5 min
Field notes from building and operating a small private GPU platform with Harvester, K3s, and a GitLab -> Flux delivery loop.
- Hybrid/On-Prem GPU: The Boring GitOps PathDec 29, 2025 · 5 min
The contracts I use to operate mixed-vendor GPU capacity with Kubernetes and Flux, including cost, storage, scheduling, and rollback limits.
- GPU Failure Modes: What Breaks and How to Debug ItDec 29, 2025 · 5 min
A cross-vendor incident workflow for separating GPU scheduling, driver, memory, model-loading, and request-path failures.
- GPU Cost Baseline: What to Measure, What LiesDec 29, 2025 · 5 min
A reproducible way to join GPU capacity, utilization, request outcomes, and cost without pretending allocation equals useful work.
- AI Infra Readiness Audit: What I Check (and What You Get)Dec 29, 2025 · 4 min
The evidence I collect before recommending changes to GPU cost, inference reliability, or deployment operations.