Hard isolation per tenant
Each guest gets a whole card. No time-slicing arbitration, no noisy-neighbour interference, and an accounting boundary that is a hardware boundary.
Inference economics come down to two numbers: how much of the machine you can hand to the model, and how fast the weights land in VRAM. Unimatrix0 optimises both. Everything else on this page follows from those two.
Design targets for a reference deployment. We publish measured numbers only after measuring them on your hardware and your model.
On a converged stack, every host reserves 16–32 GB for control-plane daemons, container runtimes and shims. On a 256 GB node serving a 70B model at long context, that reservation is the difference between fitting and not fitting.
How to use this commercially. Model memory is roughly params × bytes-per-param plus KV cache that grows with concurrency and context. Take the RAM you recover, subtract the KV cache you must reserve for your concurrency target, and the remainder is your next partial layer or your next replica. That arithmetic is the whole pitch.
256 GB node · 70B model · 8 replicas, long context CONVERGED STACK reserved for control plane │███ │24 GB guest allocator overhead │█ │6 GB model weights │███████████████████ │140 GB KV cache @ target concurrency │██████████ │72 GB ✗ does not fit — reduce concurrency or buy RAM UNIMATRIX0 Dom0 · fixed │█ │2 GB model weights │███████████████████ │140 GB KV cache @ target concurrency │██████████ │72 GB headroom / 9th replica │██████ │42 GB ✓ fits, with 42 GB spare Recovered on a 60-node estate: ~1.7 TB.
Drivers are not installed in the hypervisor. The physical device is assigned to a dedicated guest. That is what makes multi-tenant GPU economics possible — and it is why a bad model cannot take down the machine.
Each guest gets a whole card. No time-slicing arbitration, no noisy-neighbour interference, and an accounting boundary that is a hardware boundary.
Driver fault, out-of-memory in the runtime, a pathological kernel: it kills one appliance. The hypervisor, the other tenants and the traditional VMs keep serving.
Eight cards in a host means eight independent failure and capacity domains. Losing one costs you one tenant's throughput, not the node.
| Concern | Shared-driver model | Passthrough to isolated guest |
|---|---|---|
| VRAM fidelity | Virtual framebuffer overhead | 100% of the card, native addressing |
| Driver fault blast radius | Host kernel and every tenant | One appliance instance |
| Tenant boundary | Software policy | Hardware assignment |
| Runtime upgrade | Requires host maintenance | Rebuild the appliance; host untouched |
| Chargeback / metering | Approximate | Per-card, per-tenant, exact |
| Hybrid CPU inference | Mixed paths, mixed overhead | Same scheduling model for both |
In a container platform, scaling a model server means pulling tens of gigabytes over HTTP per replica — during autoscaling, exactly when your network is busiest. Unimatrix0 keeps weights on shared block storage and attaches them read-only.
40 GB weights · 8 replicas coming online HTTP pull path image layer ████████████ 2–5 min per replica + registry round trip + autoscaling burst while your WAN is saturated × 8 replicas = ~320 GB in, nothing cached on next deploy: repeat from zero Block attach path read-only attach from shared array load into VRAM ████████████ ~4.8s + one-time ingest onto the array × 8 replicas = same 40 GB on disk, 8 readers on next deploy: warm again in seconds Same bytes. Different fabric.
State it honestly to engineers. This is fast block I/O, not magic. Sustained throughput is bounded by your array's read bandwidth — we size that with you, and we will tell you if your array will become the bottleneck for your replica count.
Inference and batch evaluation requests are published to a shared queue. Each node watches its own VRAM utilisation and thermal state, and an idle GPU claims the work. This is the same mechanism that places virtual machines — one model, not two.
A batch evaluation or an agent request lands on a cluster-wide queue. You do not choose a node. Nodes are anonymous capacity.
Each node evaluates free VRAM, GPU temperature and existing load, then claims only what it can genuinely serve.
The job runs inside that card's appliance. Tokens stream back over the bus to the caller. No orchestration layer in the path.
If a card's host fails, the workload moves by lease, and the endpoint resumes from the queue rather than dropping the conversation.
| Workload shape | How Unimatrix0 runs it | What makes it efficient |
|---|---|---|
| High-concurrency chat / RAG | One card per replica, replicas spread across racks | Anti-affinity placement, per-tenant VRAM accounting, no shared-driver contention |
| Batched offline evaluation | One queued job, claimed by whichever cards free up first | Work-stealing queue replaces an autoscaler and a job controller |
| Fine-tuning / training | Long-running appliances with reserved cards | Weights and datasets already on the array; a driver fault is contained |
| Model hot-swap / canary | Attach a different weights artefact, drain the old replica | Sub-5 s transition; rollback is a reattach |
| CPU-only embeddings / small models | Same queue, same placement, no card required | One scheduling model across CPU and GPU — no second orchestrator |
| Inference for our own agents | On-site sampling capability, fully inside your perimeter | Decisions never leave the building; no third-party API in the control path |
Appliances that want help declare that they use the platform's sampling capability. When an AI appliance is present, the request is routed to a model running on your own hardware. When it is absent, those features simply do not appear — nothing breaks.
scenario: HA failover completed, cluster looks healthy platform event node pg-07 restarted on umx-7c2e (lease acquired) 23 sibling restarts in the same batch │ ▼ local model on umx-7c2e (on your hardware) │ │ reads: xen://logs · sanlock://leases · node metrics ▼ recommendation "23 restarts on one surviving host is consistent with a rack power event and a stale lease cache. Verify r12 power feed before accepting further HA placements into that domain." │ ▼ operator decision the model does not place, fence, or approve anything. A human accepts the change.
If your inference has to leave the building, you do not have private AI — you have a contract, a data-processing agreement, and a latency budget you do not control.
Prompts, completions, weights and telemetry stay on your array. There is no telemetry endpoint, no usage reporting call, no licence server.
Nodes are diskless and discover peers locally. Model weights are ingested deliberately, once. The platform works with no internet route.
Retention, deletion and residency are storage policy, not a vendor's data-handling page.
Because the interface is an open protocol, your own models and workflows integrate without a vendor partnership.
Send us the workload profile — model size, target context, peak and sustained request rate, and what the data may not do. We will come back with a sizing model and an honest answer about whether we fit.