AI Infrastructure Readiness
Before you can optimize GPU costs or ship reliable inference, you need to know where you stand. This guide covers the diagnostic, the deep dives, and the implementation paths.
The Diagnostic
A fixed-scope audit that turns assumptions into baselines and gives you an executable 90-day roadmap.
AI Infra Readiness Audit
2-3 weeks. Scorecard, cost model, risk register, and prioritized roadmap.
The audit covers GPU cost baselines, reliability gaps, deployment safety, and failure modes. You get a written assessment and a sequenced plan with owners.
Deep Dives
Technical guides covering specific aspects of GPU infrastructure: cost, reliability, architecture, and debugging.
3 min read
AI Infra Readiness Audit: What I Check (and What You Get)
A practical checklist for auditing production AI infrastructure: GPU cost baselines, reliability risks, and an executable roadmap.
4 min read
GPU Cost Baseline: What to Measure, What Lies
Before you can cut GPU costs, you need to measure them correctly. Here is what to track and what the cloud console will not tell you.
5 min read
GPU Failure Modes: What Breaks and How to Debug It
Common GPU infrastructure failures in production and how to diagnose them before they become incidents.
4 min read
Hybrid/On-Prem GPU: The Boring GitOps Path
A practical guide to running GPU workloads on-prem or hybrid, using Kubernetes and GitOps patterns that make operations boring.
6 min read
SLOs for Inference: Latency, Errors, Saturation
How to define meaningful SLOs for production inference workloads, and what to do when they break.
Related Resources
Additional posts on GPU infrastructure, Kubernetes, and inference workloads.
10 min read
Five Voices, 200K Tokens, One Consumer GPU: Giving Simulated Minds a Whole Life to Remember
We rebuilt our Jungian multi-agent lab so each archetype carries its entire life — every situation, everything heard, everything said — as an append-only journal that fills a 228K-token context window on a 24 GB Radeon. Prefix caching makes depth nearly free; when the window fills, the agents dream their oldest memories into a durable core. Here is the architecture, the numbers, and what broke along the way.
4 min read
Gaming Mode: Turning an Inference GPU Node into a Moonlight Host — Declaratively
FlexInfer had a "gaming mode" that had never once run a game. I killed the stub, proved the substrate with a 30-minute kill-test, and rebuilt it as a declarative GamingSession CR that drains inference and streams a GPU-accelerated game over Moonlight.
10 min read
State of the Platform, July 2026
The quarterly snapshot of the whole stack: FlexInfer went multimodal, Mills merged its first autonomous work, a flight recorder appeared, and the dashboards grew a control plane. What shipped, what the numbers say, and what is deliberately not done.