The archive, mapped
Every blog post and case study on the site with the principles and playbooks each piece argues for. Filter by any principle or playbook to see just the evidence behind it.
Matching the filter
5 itemsHow fi-fhir separates feed-specific variance from parsing, semantic events, durable delivery, and workflow policy—and where the pre-1.0 platform still has work to do.
- Deploying MLC-LLM on Dual RX 7900 XTX GPUs: Debugging VRAM, KV Cache, and K8s GPU SchedulingJan 4, 2026 · 13 min
What actually broke when I deployed MLC-LLM across two RX 7900 XTX nodes, and the fixes that made it stable: quantization, KV cache sizing, and Kubernetes GPU hygiene.
- GPU Failure Modes: What Breaks and How to Debug ItDec 29, 2025 · 5 min
A cross-vendor incident workflow for separating GPU scheduling, driver, memory, model-loading, and request-path failures.
A de-identified field pattern for containing patient-match risk when filter and mode parameters interact in undocumented ways.
- Running LLMs on Radeon GPUs with ROCmNov 20, 2025 · 5 min
What still works, what changed, and which guardrails matter when you run AMD Radeon GPUs for always-on inference.