The archive, mapped
Every blog post and case study on the site with the principles and playbooks each piece argues for. Filter by any principle or playbook to see just the evidence behind it.
All content
34 itemsI put fi-fhir and edilint through Mills for the first time. It merged work I did not write, then exposed a test gate passing on the wrong tree and fixed it against itself.
- The Agent Gets Read-Only Credentials: A Trust Ladder for AI in Regulated SystemsAug 29, 2026 · 4 min
A design pattern for supervised AI investigations: scoped read access, reviewed drafts, and failure tests before expanding permissions.
- The Factory Upgrades Its Own Foundation: An MCP Spec MarathonAug 29, 2026 · 6 min
MCP shipped a breaking specification revision. I wrote the gap analysis, filed it as factory backlog, and Mills started rebuilding the SDK it runs on.
- A Green Run That Checks Nothing: Seven Silence Modes, Seventy Repos, One DayAug 16, 2026 · 16 min
A fleet-wide maintenance day exposed seven ways CI can report success without useful work—and the merge-queue discipline and recurring checks built to catch them.
- The File That Failed With Zero Errors: A Field Guide to Invisible EDI DefectsAug 14, 2026 · 5 min
Four defect classes that healthcare interchange parsers can accept silently, with synthetic samples, real linter output, and the pre-send gate that catches them.
- Twin Life on a Doubled Pool: One Day of Canary-Driven Inference EngineeringJul 26, 2026 · 12 min
A twin-lane canary promoted prefix caching, FP8 KV cache, and a larger context pool while catching five regressions before they reached the primary workload.
- Five Voices, 200K Tokens, One Consumer GPU: Giving Simulated Minds a Whole Life to RememberJul 24, 2026 · 13 min
How five simulated minds share one 24 GB Radeon, keep 200K-token life histories, reuse them through prefix caching, and compact old memories without erasing continuity.
- State of the Platform, July 2026Jul 1, 2026 · 10 min
A July 2026 platform snapshot: multimodal inference, Mills delivery, flight recording, and the operational gaps that remained after the demos worked.
I replaced FlexInfer’s unused gaming-mode stub with a declarative GamingSession resource that drains inference and turns a GPU node into a Moonlight host.
- Finding the Real Context Ceiling: Needle-Benchmarking Forced RoPE ExtrapolationJun 25, 2026 · 5 min
Needle-retrieval tests on a 35B model passed near 63,000 prompt tokens and failed at longer contexts, even when vLLM loaded without a memory error.
- Loom Mills: From Agent Swarms to Software Production LinesMay 2, 2026 · 7 min
The May 2026 design for Loom Mills: turn roadmap intent into scoped work, verify each change, and keep costs, failures, and operator controls visible.
- A One-Page AI Usage Policy That Actually WorksApr 20, 2026 · 4 min
A short, adoptable AI usage policy for engineering teams: what to put on the page, what to leave off, and why the policy matters less than the habits it makes explicit.
- The First 90 Days: Introducing AI-Assisted Dev to a New TeamApr 20, 2026 · 8 min
A 90-day rollout for AI-assisted development: start with bounded habits, measure the work, and standardize only after the team has evidence.
- Getting Gemma 4 Running on a Radeon 7900 XTX (with and without TurboQuant)Apr 4, 2026 · 8 min
Field notes from serving Gemma 4 E4B on Radeon: the stable TRITON path, an experimental TurboQuant long-context lane, and the GPTQ work still in progress.
- Build Your Own Legs Before the Crutches FailMar 9, 2026 · 13 min
How I use AI drafts to build engineering judgment through inspection, regression tests, code review, and repeatable release checks.
- Repo Design Patterns for AI-Assisted Dev: Control Loops, Hooks, and MemoryFeb 9, 2026 · 6 min
Treat the repository as a control system: versioned instructions, explicit workflows, fast hooks, durable context, and visible failure states.
- Loom: One Registry, Many AI Coding AssistantsFeb 9, 2026 · 7 min
How Loom generates and checks MCP configuration across nine client profiles from one canonical registry.
- Loom Mode: One Proxy, One Daemon, Faster MCPFeb 9, 2026 · 7 min
How Loom keeps MCP tooling predictable with one client entrypoint, daemon-owned process lifecycles, bounded output, and coarse-grained tools.
- Two-Lane Text GPU Allocation: Quality + Vision/Fast (Plus a Media Lane)Feb 9, 2026 · 11 min
A February 2026 GPU allocation experiment: separating text workloads, testing shared-group priorities, and keeping client routes stable as models moved.
How I evolved a small AMD GPU cluster into a CRD-driven inference platform with an OpenAI-compatible boundary, evidence-based model promotion, and safe rollouts.
How fi-fhir separates feed-specific variance from parsing, semantic events, durable delivery, and workflow policy—and where the pre-1.0 platform still has work to do.
- Deploying MLC-LLM on Dual RX 7900 XTX GPUs: Debugging VRAM, KV Cache, and K8s GPU SchedulingJan 4, 2026 · 13 min
What actually broke when I deployed MLC-LLM across two RX 7900 XTX nodes, and the fixes that made it stable: quantization, KV cache sizing, and Kubernetes GPU hygiene.
- SLOs for Inference: Latency, Errors, SaturationDec 29, 2025 · 5 min
How I define request, cold-start, queue, and model-quality objectives for an inference service without confusing causes with user outcomes.
- Standing Up a GPU-Ready Private AI Platform (Harvester + K3s + Flux + GitLab)Dec 29, 2025 · 5 min
Field notes from building and operating a small private GPU platform with Harvester, K3s, and a GitLab -> Flux delivery loop.
- Hybrid/On-Prem GPU: The Boring GitOps PathDec 29, 2025 · 5 min
The contracts I use to operate mixed-vendor GPU capacity with Kubernetes and Flux, including cost, storage, scheduling, and rollback limits.
- GPU Failure Modes: What Breaks and How to Debug ItDec 29, 2025 · 5 min
A cross-vendor incident workflow for separating GPU scheduling, driver, memory, model-loading, and request-path failures.
- GPU Cost Baseline: What to Measure, What LiesDec 29, 2025 · 5 min
A reproducible way to join GPU capacity, utilization, request outcomes, and cost without pretending allocation equals useful work.
- AI Infra Readiness Audit: What I Check (and What You Get)Dec 29, 2025 · 4 min
The evidence I collect before recommending changes to GPU cost, inference reliability, or deployment operations.
- Optimizing Real-Time Kubernetes VisualizationsDec 25, 2025 · 8 min
What the current FlexDeck code does to keep Canvas 2D and Three.js cluster views responsive, and how its performance harness avoids unsupported frame-rate claims.
A de-identified field pattern for containing patient-match risk when filter and mode parameters interact in undocumented ways.
- Welcome to My HomelabNov 27, 2025 · 6 min
An August 2026 tour of the Harvester, K3s, GitLab, Harbor, Flux, and mixed-GPU platform behind the systems I publish here.
- Running LLMs on Radeon GPUs with ROCmNov 20, 2025 · 5 min
What still works, what changed, and which guardrails matter when you run AMD Radeon GPUs for always-on inference.
- Building Practical AI AgentsNov 15, 2025 · 5 min
A source-backed guide to agent reliability: explicit state, bounded tools, verification gates, and visible failure.
How FlexDeck combines Kubernetes, Flux, CI, observability, and model state without hiding freshness, access boundaries, or the limits of a homelab control surface.