AI Infrastructure Readiness
Before you can optimize GPU costs or ship reliable inference, you need to know where you stand. This guide covers the diagnostic, the deep dives, and the implementation paths.
The Diagnostic
A fixed-scope audit that turns assumptions into baselines and gives you an executable 90-day roadmap.
AI Infra Readiness Audit
2-3 weeks. Scorecard, cost model, risk register, and prioritized roadmap.
The audit covers GPU cost baselines, reliability gaps, deployment safety, and failure modes. You get a written assessment and a sequenced plan with owners.
Deep Dives
Technical guides covering specific aspects of GPU infrastructure: cost, reliability, architecture, and debugging.
4 min read
AI Infra Readiness Audit: What I Check (and What You Get)
The evidence I collect before recommending changes to GPU cost, inference reliability, or deployment operations.
5 min read
GPU Cost Baseline: What to Measure, What Lies
A reproducible way to join GPU capacity, utilization, request outcomes, and cost without pretending allocation equals useful work.
5 min read
GPU Failure Modes: What Breaks and How to Debug It
A cross-vendor incident workflow for separating GPU scheduling, driver, memory, model-loading, and request-path failures.
5 min read
Hybrid/On-Prem GPU: The Boring GitOps Path
The contracts I use to operate mixed-vendor GPU capacity with Kubernetes and Flux, including cost, storage, scheduling, and rollback limits.
5 min read
SLOs for Inference: Latency, Errors, Saturation
How I define request, cold-start, queue, and model-quality objectives for an inference service without confusing causes with user outcomes.
Related Resources
Additional posts on GPU infrastructure, Kubernetes, and inference workloads.
14 min read
Five Voices, 200K Tokens, One Consumer GPU: Giving Simulated Minds a Whole Life to Remember
How five simulated minds share one 24 GB Radeon, keep 200K-token life histories, reuse them through prefix caching, and compact old memories without erasing continuity.
4 min read
Gaming Mode: Turning an Inference GPU Node into a Moonlight Host — Declaratively
I replaced FlexInfer’s unused gaming-mode stub with a declarative GamingSession resource that drains inference and turns a GPU node into a Moonlight host.
10 min read
State of the Platform, July 2026
A July 2026 platform snapshot: multimodal inference, Mills delivery, flight recording, and the operational gaps that remained after the demos worked.