Skip to main content
Self-hosted AI operations

FlexInferPrivate AI, operated like real infrastructure

Run and customize open models on your own GPUs, behind an OpenAI-compatible gateway. Loom governs how agents reach tools and context, MentatLab turns model calls into observable DAG workflows, and fi-fhir proves the stack on real healthcare data.

OpenAI-compatible APIGPU-aware, scale-to-zeroMCP-governed agent access
Route preview
Private runtime boundary
Ready
Boundary
Private
Ops
GitOps
Routes
4 lanes
Runtime
GPU-aware
Gateway
OpenAI-compatible
Context
Bounded MCP
Workflow
DAG-visible
flexinfer.route.yaml
model: gemma4-26b-a4b
route: /v1/chat/completions
placement: gpu-pool/ready
policy: tools.allowed[3]
Measured on the clustersnapshot

Verified FlexInfer benchmark results

Serving throughput across language, embedding, and image workloads. The live feed takes over when available; otherwise this shows the latest verified cluster snapshot.

Snapshot verified

106.0tok/s recorded peak
FlexInfer docs

Snapshot scope. These legacy benchmark-result records preserve backend, hardware class, memory, and throughput, but not prompt, batch, or sample counts. Treat them as operational baselines, not a cross-vendor leaderboard.

Inspect the source records
Platform posture

Four products, one operating boundary

The stack is intentionally layered: run and customize models inside your boundary, govern how agents reach tools and context, orchestrate repeatable DAG workflows over direct API models, and use fi-fhir when the workload is sensitive healthcare ETL instead of generic demo data.

Integration points

How platform surfaces connect in production

Use these contracts to map deployment and integration boundaries before implementation.

Core Stack

Loom suite, FlexInfer, fi-fhir, and MentatLab

FlexInfer runs the models, Loom governs context and policy, MentatLab gives operators the DAG view, and fi-fhir carries the healthcare workload. Product pages explain where each piece fits; docs and playground show it running.

Operational surface

MentatLab mission control

Loom Core governs context routing and policy boundaries. MentatLab provides the DAG design and run-visibility layer over direct API model calls, internal services, and private FlexInfer endpoints.

Mobile

Loom Companion

SwiftUI app for fleet monitoring, session management, real-time alerts, and lightweight operator control from iPhone and iPad.

Enterprise capabilities

Operator controls already shipping across the platform stack

MCP gateway routing, RBAC, sandbox execution, and OTel-traced operations are available now. HUD cost and audit visibility are landing next.

In progress

HUD Cost Dashboard
In progress

Cost monitoring integration: loom/cost-stats RPC, CostMonitor polling, SSE events, and OverviewPanel KPI tile.

HUD RBAC + Audit Visibility
In progress

RBAC config RPC, denied-calls ring buffer, ServersPanel RBAC sub-tab, and OverviewPanel badge.

Available now

MCP Gateway
Available

Centralized MCP routing via loom proxy with streamable HTTP transport, bearer/OIDC/mTLS auth, and hub failover.

Single context ingress with controlled routing, auditing, and automatic local fallback.

Role-based Access Control (RBAC)
Available

Role-aware permissions for MCP tool access with audit trail, cost tracking, and OAuth 2.1.

Enforces least-privilege context access with auditable decision logs across teams and environments.

Sandbox Executor (Docker + K8s)
Available

mcp-devbox sandbox runtime with Docker and Kubernetes backends for isolated agent execution.

Runs builds, tests, and automation in controlled containers with consistent isolation and audit trails.

Operational Foundations
Available

OTel tracing across all 59 MCP servers, JSON log correlation, observability stack, and deployment controls.

Full production observability with distributed tracing, structured logging, and repeatable deployment workflows.

Start Here

Pick a product, then prove it in the playground

Docs, playground, and product pages stay aligned on the same config shapes and schemas, so what you validate here is what you deploy.

Consulting

Bring it into production

If you want this stack inside your environment, the fastest path is a scoped audit or build engagement.