Agent-ready via MCP No control-plane database GPU passthrough

Infrastructure that thinks in your data centre.

Unimatrix0 is a stateless, masterless private cloud for bare metal. It runs your traditional virtual machines, hosts your GPU inference at near-100% memory efficiency, and exposes the entire cluster to AI agents through the Model Context Protocol — on one fabric, with no database to lose.

Boot a node, and it is production. No installer, no controller to deploy, no control plane to lose.

Built on components your team already knows
Xen Alpine Linux NATS Sanlock Model Context Protocol NFS · iSCSI · Fibre Channel · NVMe-oF CUDA · ROCm vLLM · Ollama · PyTorch · ONNX Claude Desktop · Cursor · LangChain
The problem

The control plane is the single point of failure.

Every mainstream platform puts a stateful database — SQL, etcd, corosync — between you and your hardware. It is the thing that must be patched, backed up, quorum-tuned, and upgraded in a specific order. It is the reason upgrades are events, and why the smallest VM cannot be scheduled if the database is sad.

State that can go stale

A control database is a second, always-imperfect copy of reality. It drifts from the hypervisor, it blocks on quorum, and it turns a partial network partition into a cluster-wide outage rather than a local one.

Reserved capacity you cannot use

Converged stacks reserve 16–32 GB per host for control-plane daemons, container runtimes and shims. On a 64-core AI node that is memory you bought, paid for, and cannot hand to a model.

Automation that arrives late

When your orchestrator can be driven by an API, but not understood by an agent, every capacity decision still routes through a human reading a dashboard at 3 a.m.

The inversion

Five decisions that change everything downstream.

Unimatrix0 does not improve the traditional control plane. It removes it and derives the same outcomes from the hypervisor itself.

01 / SOURCE OF TRUTH

The hypervisor and the storage array are the database.

Live workload state is read from the hypervisor toolstack. Configuration lives as plain files on shared storage, next to the disks it describes. There is no cache to invalidate, no migration to run, and no backup of the control plane — because there is nothing to back up that is not already on your array.

02 / SYMMETRY

Every node is identical, and any node can serve.

Nodes are peers, not servers and clients. Any node accepts an API call, publishes work, or claims it. There is no leader to elect, no management role, and no node whose loss degrades the cluster.

03 / IMMUTABILITY

The operating system is a read-only image in RAM.

Nodes boot a squashed filesystem that never touches local disk and cannot be modified at runtime. A node is disposable: power-cycle it and it is byte-identical. A compromised node cannot persist, and a fleet-wide upgrade is a reboot from new media.

04 / LEAST PRIVILEGE

The hypervisor carries no workloads.

Networking, storage, backup, AI runtimes and protocol parsers all run in isolated guests with their own PCI devices. A kernel panic in a CUDA driver takes down one model server, not the host.

05 / PROTOCOL-FIRST AUTOMATION

Every operation is a standard MCP tool, not a private endpoint.

Provision a VM, drain a node, read a lease table, snapshot a workload — each is exposed through the open Model Context Protocol with the same permissions, approvals and audit trail as a human clicking in the console. Every MCP client already speaks it. No SDK, no plugin, no integration project.

Side by side

What the comparison table actually looks like.

Figures below are design targets for a reference deployment on typical enterprise hardware. They describe the architectural ceiling, not a benchmark of your workload.

Comparison of Unimatrix0 against Kubernetes with KubeVirt and commercial hypervisor stacks
Dimension Kubernetes + KubeVirt Commercial hypervisor Unimatrix0
Control-plane state Stateful etcd cluster with Raft quorum Central SQL database, vCenter None. Hypervisor + storage are the only state
RAM reserved per host 16–32 GB 4–8 GB ~2 GB fixed Dom0 — the rest is workload capacity
Usable host memory Reduced; reserved by daemons Reduced; reserved by the stack >99% to workloads and model context
VRAM available to models Partial — shared with host drivers Partial — vGPU layer overhead 100% via hardware passthrough
Node cold boot 3–5 min 2–5 min 18–25 s from live RAM
Failure detection Seconds; depends on health probes Heartbeat + storage heartbeats ~2 s signal, then disk-level fencing
Split-brain protection Raft quorum Storage heartbeats Atomic disk leases — two hosts can never own one disk
40 GB model load 2–5 min over HTTP Moderate <5 s — SAN block mount into VRAM
Agent / AI interface CRDs + bespoke controllers Proprietary API Standard MCP — tools, resources, prompts
GPU driver isolation Host kernel Host / vGPU layer Isolated guest with the physical card
Upgrade mechanism Rolling component upgrades, ordering-sensitive Rolling vCenter + host patches Drain, reboot from new image, rejoin
Persistence after compromise Attacker can persist on the host Attacker can persist on the host None. Reboot eradicates runtime changes

Reference deployment: 3 nodes, dual-socket Xeon, 256 GB RAM, 2 GPUs per node, 10 GbE iSCSI SAN. Measured behaviour varies with hardware and workload; we publish our own numbers only after we have measured them on your profile.

AI first

Every megabyte of RAM you did not spend on a control plane is context window.

Inference economics are decided by two numbers: how much of the machine you can give to the model, and how fast you can get weights into VRAM. Unimatrix0 is engineered around exactly those two numbers.

  • ~2 GB fixed overhead. A fixed, small Dom0 on every host — not a scaling control plane.
  • 100% of VRAM, one card per tenant. Physical passthrough, not shared virtual frames.
  • Weights already on the array. Models live on shared block storage and attach read-only.
  • The cluster is the scheduler. Idle GPUs claim queued inference work by themselves.
         SHARED BLOCK STORAGE  (SAN / NVMe-oF)
  model weights .safetensors .gguf · read-only attach
  ────────────────────────────────────────────────────
         │ block attach          │ block attach
               ▼                       ▼
┌─ umx-node-01          ─┐    ┌─ umx-node-04          ─┐
│ Dom0  2GB  99% free    │    │ Dom0  2GB  99% free    │
│ sys-ai  H100 #0   71%  │    │ sys-ai  H100 #0   12%  │
│ sys-ai  H100 #1   88%  │    │ sys-ai  A100 #0   31%  │
└────────────────────────┘    └────────────────────────┘
           │                            │
           │  tasks.ai.global  (queue)  │
└──────────────────────────┬───────────────────────────┘
                           ▼

         node-04 claims the job — VRAM is free.
                 No scheduler. No YAML.
MCP fabric

Your infrastructure already speaks the agent protocol.

Point any MCP client at a node. It discovers tools, reads live state, and runs operations — under the same policy engine that governs the web console and the CLI.

Tools

unimatrix_create_vm, unimatrix_drain_node, ai_assign_workload. Annotated read-only or destructive, so clients know what they are about to do.

Resources

xen://node/domains, sanlock://leases, gpu://inventory. Live cluster state as addressable URIs an agent can read.

Prompts

Reusable procedures shipped with the product: root-cause a failed failover, plan a restore, rebalance placement — as structured prompts, not tribal knowledge.

Guardrails

Read scopes run freely. Destructive calls raise an approval a human accepts. Every call is attributed, validated against a schema, and written to an audit log.

2GB
Per host, fixed
Total control-plane overhead. It does not grow with fleet size.
18–25 s
Node cold boot
Firmware to a schedulable hypervisor, from RAM.
2s
Failure detection
Then a disk lease decides who may restart the workload.
99%+
Host RAM to workloads
No reserved pool for the machinery running the machinery.
Scale

One node to one rack to one continent — same product.

Growth is a topology problem, and topology is explicit. Nodes carry site, zone and rack labels. Placement spreads replicas across fault domains, reserves capacity for the worst plausible loss, and refuses to take on work it could not survive.

  • Anti-affinity that is real. Three database replicas land in three racks, not three sockets.
  • N+1 reserve. The console answers "can we survive losing rack R?" before you find out.
  • Drain by domain. Evacuate an entire rack with rate control, not node by node.
  • Many clusters beat one big cluster. Past a few hundred nodes, more failure domains wins.
site: eu-west-1
├── cluster eu-west-1a  — primary
│   ├── rack r11   node-01  node-02  ♥ leader
│   ├── rack r12   node-03  node-04
│   └── rack r13   node-05  node-06  (drained)
│
│   placement: pg-primary → anti-affinity: rack
│   reserve:    N+1 rack ................. satisfied
│   lose rack r12 → restart 48 VMs?  yes, 62s
│   lose rack r13 → restart 48 VMs?  no — insufficient headroom
│
└── cluster eu-west-1b  — recovery
    connected by WAN gateway · runs autonomously if the link drops
Get started

Bring us the workload you cannot get to run.

Tell us the model, the concurrency, the recovery objective, and the constraints — air-gapped, sovereign, regulated. We will tell you honestly whether we are the right fit, and what a pilot would look like on your hardware.

No signup wall. You will speak to the engineers who build the platform.