Domain zero
A fixed allocation — the same on a 32-core node and a 128-core node. No growing control plane.
A node is an immutable Alpine Linux image running the Xen hypervisor in a tiny domain zero. Nodes share a storage array and a message bus. That is the entire platform — and it is enough to run enterprise virtualisation, GPU inference, and an agent-driven MCP fabric at the same time.
Every layer below is optional in the sense that nothing depends on a higher layer existing. Remove any box and the rest keeps running.
┌──────────────────────────────────────────────────────────────────────────┐ │ SHARED STORAGE ARRAY SAN · NAS · NVMe-oF · local+replicated │ │ VM definitions · disks · images · ISO library · disk leases │ └──────────────────────────────────────────────────────────────────────────┘ │ block I/O │ block I/O ┌─ node-a1 ─┐ ┌─ node-a2 ─┐ ┌─ node-a3 ─┐ │ ┌───────────────────────┐ │ │ ┌───────────────────────┐ │ │ ┌───────────────────────┐ │ │ │ ┌───────────────────┐ │ │ │ │ ┌───────────────────┐ │ │ │ │ ┌───────────────────┐ │ │ │ │ │ console / UI │ │ │ │ │ │ console / UI │ │ │ │ │ │ console / UI │ │ │ │ │ └───────────────────┘ │ │ │ │ └───────────────────┘ │ │ │ │ └───────────────────┘ │ │ │ │ ┌───────────────────┐ │ │ │ │ ┌───────────────────┐ │ │ │ │ ┌───────────────────┐ │ │ │ │ │ daemon api · mcp │ │ │ │ │ │ daemon api · mcp │ │ │ │ │ │ daemon api · mcp │ │ │ │ │ └───────────────────┘ │ │ │ │ └───────────────────┘ │ │ │ │ └───────────────────┘ │ │ │ │ ┌───────────────────┐ │ │ │ │ ┌───────────────────┐ │ │ │ │ ┌───────────────────┐ │ │ │ │ │ Dom0 · Xen xl │ │ │ │ │ │ Dom0 · Xen xl │ │ │ │ │ │ Dom0 · Xen xl │ │ │ │ │ └───────────────────┘ │ │ │ │ └───────────────────┘ │ │ │ │ └───────────────────┘ │ │ │ │ ┌───────────────────┐ │ │ │ │ ┌───────────────────┐ │ │ │ │ ┌───────────────────┐ │ │ │ │ │ sys-* appliances │ │ │ │ │ │ sys-* appliances │ │ │ │ │ │ sys-* appliances │ │ │ │ │ └───────────────────┘ │ │ │ │ └───────────────────┘ │ │ │ │ └───────────────────┘ │ │ │ └───────────────────────┘ │ │ └───────────────────────┘ │ │ └───────────────────────┘ │ └─────────────────────────────┘ └─────────────────────────────┘ └─────────────────────────────┘ MESSAGE BUS + PEER MESH peer discovery · task queues · events · HA signals ────────────────────────────────────────────────────────────────────────── Appliances run as isolated guests. None of them live in the hypervisor. sys-ai sys-net sys-backup sys-mcp-gateway GPU inference routed networks snapshots + external agent vLLM · Ollama BGP · WireGuard offsite ship exposure PyTorch · ONNX firewall (passthrough) rate limiting
Domain zero is a memory-resident Linux with the hypervisor and one daemon. It is deliberately too small to be interesting, which is the point: it is identical everywhere, so there is no fleet skew and nothing to patch per host.
A fixed allocation — the same on a 32-core node and a 128-core node. No growing control plane.
Everything above that line is memory and VRAM your workloads and model context can use.
Nodes boot from read-only media and keep nothing. Power-cycle returns a clean machine.
One per node serves the API, the console, the MCP endpoint and the message bus.
Why this matters commercially. A fixed 2 GB cost per host means the overhead of running your private cloud does not grow with the fleet. On a 60-node estate that is roughly 1.7 TB of memory returned to workloads — and about 1.7 TB of hardware you may not have to buy.
Split-brain is not a quorum problem you can tune away. Unimatrix0 answers it with an atomic claim on shared storage — the one thing that cannot be decided twice.
T+0.0s node-a1 loses power mid-transaction ── SIGNAL PLANE ─────────────────────── T+2.0s peers notice the missing heartbeat "node-a1 suspected down" T+2.3s surviving nodes apply randomised jitter 100–500ms, so they do not collide ── SAFETY PLANE ─────────────────────── T+2.6s node-a2 attempts the disk lease for pg-07 ✓ acquired → permitted to start ✗ denied → stands down, no-op T+6.0s node-a1 is not powering off, it is hung: hardware watchdog trips, host resets. Fencing happens BEFORE the lease expires. T+7.0s lease is genuinely free — workload restarts on a healthy node, storage re-attached. Result: exactly one node ever owns the disk. A slow network can delay recovery. It cannot produce two writers.
Quorum systems trade availability for consistency: during a partition, some nodes stop accepting writes. Lease systems on shared storage do the opposite — they stay available everywhere except on the one disk in dispute, and the dispute resolves because the storage array is the arbiter.
Be precise with your customers. We offer fast detection and strong safety. We do not offer the impossible: a lease can only be granted when the storage layer is reachable and the previous holder has stopped renewing. If your storage array lies about durability, no software above it can save you. Ask us about failure drills — we run them and we will show you the timings.
Networking, backup, AI runtimes and protocol gateways run in their own guests. They declare what they can do once, and the platform generates the API, the CLI command, the console screen and the agent tool from that single declaration.
Layer-3 routing, BGP, firewalling, VPN and high-speed fabric networking — with its own passthrough NIC, out of the hypervisor's way.
Snapshots, schedules and restores. Streams changed blocks straight off shared storage to offsite object storage without staging a copy anywhere.
GPU-attached inference guests running vLLM, Ollama, PyTorch or ONNX. A driver fault or a runaway model is contained to one appliance instance.
Optional hardened endpoint for exposing the MCP fabric beyond the management network, with rate limiting and tenant policy.
Block snapshots, pool inspection, capacity and object-store targets for images and model weights.
Wrap an existing protocol server — a network controller, a storage vendor's tooling — and it inherits the platform's names, permissions, audit and approvals.
| Declared once… | …becomes four surfaces |
|---|---|
| Operation + JSON schema | REST endpoint · MCP tool · CLI command · console form |
| Read-only / destructive flag | Client hints · inline confirmation · required approval · dry-run guard |
| Resource URI | Addressable cluster state, subscribable for live updates |
| Prompt | A reusable operating procedure an agent can invoke |
Appliances are signed and versioned. An extension incompatible with your running release is hidden on that node and reported as such — it never half-loads.
Drain the node, reboot from new media, watch it rejoin. Repeat. Mixed versions during the roll are expected and visible; each node reports exactly which release it runs.
Take a host out of service without a ticket. Workloads live-migrate off it with rate control; new placements avoid it; the console shows progress.
A new hypervisor is: enable virtualisation in firmware, boot the image, point it at the array. It discovers its peers, joins, and accepts production workloads.
Network boot. Because nodes hold no state, they can boot over the network instead of from attached media. A single release pointer becomes a fleet-wide upgrade — flip it, roll the nodes, roll back by flipping it back.
Bring the constraints: your storage, your network, your recovery objectives. We will map them to the architecture and tell you where the edges are.