Independent autonomy
Each site runs a full cluster. If the WAN drops, that site is completely unaffected. No split-brain to argue about, because the domains never shared a quorum.
One node behaves like one node. So does a hundred. What changes with size is not whether the fabric works — it is how much correlated failure it must survive, and whether the platform knows about it before you do.
We publish these because a platform that hides its limits until after the contract is signed is not worth evaluating. Workload density of roughly 40 per host is our planning figure — typical across the platforms we are compared against, and where the per-host costs are sized.
| Tier | Nodes | Running workloads | What changes, and what it needs |
|---|---|---|---|
| Evaluation | 1–16 | ~640 | Nothing. The base architecture is already sufficient at this size. |
| Standard | up to ~64 | ~2,500 | Slimmer heartbeats, event-driven failover instead of polling, a single per-node domain watcher, and an explicit small quorum core for configuration. |
| Large | up to ~250 | ~10,000 | Hierarchical topology instead of a flat mesh, sharded failover ownership so no single coordinator carries the estate, server-side paging in the console. |
| Distributed | many clusters | — | One cluster per site or failure domain, linked by gateways into a federated read-only view. Each site runs autonomously if the WAN drops. |
Past roughly 100–250 nodes, more clusters beats one bigger cluster. A single cluster is a single consensus and fencing domain. Splitting by failure domain gives you isolation, independent upgrade cadence, and a smaller blast radius. We will tell you this even though it means more deployments.
At 40 workloads per host, a rack of eight hosts holds around 320. That is the largest correlated event any single cluster must survive — and the one most platforms size for last.
What we will not claim. We will not quote you a failover time without a measurement on your hardware, your storage latency and your dependency graph. Bring us your RTO and we will run the drill, then show you the timings — including the ones we are not proud of.
restart fan-out after correlated loss 1 node lost ██████████████ 40 restarts bounded by lease release 1 rack lost ████████████████████████████████████████ 320 restarts → batched, priority-ordered, rate-limited to protect shared storage 1 site lost ████████████████████████████████████████████████████ → not a restart problem. second cluster + explicit failover. ─ capacity check the platform performs ────────── lose rack r12 → restart 48 VMs? yes · 62s lose rack r13 → restart 48 VMs? no · insufficient headroom new HA placement that would break the reserve is refused at admission.
Placement that only knows "this host has more free memory" will happily put all three replicas of your database on one rack. Topology labels make location a first-class property of every node.
THE LABEL HIERARCHY site └── cluster └── zone room · power feed · cooling └── rack └── node set per node, carried in every heartbeat, visible to the console and to agents EXTRA LABELS — same mechanism gpu=h100 nvme=true storage=local-fast used as constraints and selectors by placement policy and by agents UNLABELLED NODES default to their own rack → degrades gracefully to no topology awareness
Sites are never part of the same cluster. WAN latency and partitions break storage-lease timing and consensus timing — pretending otherwise is how multi-site hypervisors lose data. Sites are linked instead.
Each site runs a full cluster. If the WAN drops, that site is completely unaffected. No split-brain to argue about, because the domains never shared a quorum.
A single console across sites, read-only across boundaries. See capacity, inventory and events everywhere without pretending it is one coherent cluster.
Asynchronous disk replication between sites, plus configuration checkpoints. Site failover is an explicit, operator-approved action with a runbook — not something a flapping link triggers at 03:00.
Each site mirrors the image and image library locally, so provisioning and upgrades do not depend on a distant link — and low-cost sites need no WAN bandwidth.
Honest boundary. There is no automatic cross-site high availability, and there will not be. A cluster never restarts another cluster's workloads, because they do not share a fencing domain. Disaster recovery is an operation you perform deliberately.
Alert on a shrunken core, replica lag and configuration checkpoint age. Those are the conditions that precede a cold start.
Keep checkpoints off-cluster and restore from them on a schedule. An untested restore path is a guess.
Distinct cluster tokens for administrative and node traffic. A shared secret means a single leak is a cluster-wide one.
Ship the message bus with the platform and upgrade it through the same rolling reboot. Finish rollouts promptly — mixed images are expected, not indefinite.
Nodes, workload density, storage throughput and control-plane bandwidth. The last one is small — keep it that way, and keep migration traffic off the management path.
Node loss, domain loss, consensus member loss, storage latency spikes, cold start. Measure recovery times and record them. Repeat quarterly.
Node counts today and in three years, where the racks and power feeds are, what your recovery objectives are, and which workloads cannot leave the site. We will model the topology and tell you where the fault domains should be.