Scale

Growth is a topology problem.

One node behaves like one node. So does a hundred. What changes with size is not whether the fabric works — it is how much correlated failure it must survive, and whether the platform knows about it before you do.

Capacity tiers

What we support, and what we will tell you not to do.

We publish these because a platform that hides its limits until after the contract is signed is not worth evaluating. Workload density of roughly 40 per host is our planning figure — typical across the platforms we are compared against, and where the per-host costs are sized.

Tier Nodes Running workloads What changes, and what it needs
Evaluation 1–16 ~640 Nothing. The base architecture is already sufficient at this size.
Standard up to ~64 ~2,500 Slimmer heartbeats, event-driven failover instead of polling, a single per-node domain watcher, and an explicit small quorum core for configuration.
Large up to ~250 ~10,000 Hierarchical topology instead of a flat mesh, sharded failover ownership so no single coordinator carries the estate, server-side paging in the console.
Distributed many clusters — One cluster per site or failure domain, linked by gateways into a federated read-only view. Each site runs autonomously if the WAN drops.

Past roughly 100–250 nodes, more clusters beats one bigger cluster. A single cluster is a single consensus and fencing domain. Splitting by failure domain gives you isolation, independent upgrade cadence, and a smaller blast radius. We will tell you this even though it means more deployments.

The arithmetic nobody shows you

A rack is not eight independent events.

At 40 workloads per host, a rack of eight hosts holds around 320. That is the largest correlated event any single cluster must survive — and the one most platforms size for last.

  • Node loss: up to 40 restarts, target — bounded by disk lease release.
  • Rack loss: up to 320 restarts. Restarted together, in dependency order, rate-limited.
  • Site loss: not a restart problem. This is where a second cluster earns its cost.

What we will not claim. We will not quote you a failover time without a measurement on your hardware, your storage latency and your dependency graph. Bring us your RTO and we will run the drill, then show you the timings — including the ones we are not proud of.

restart fan-out after correlated loss

1 node lost
  ██████████████ 40 restarts   bounded by lease release

1 rack lost
  ████████████████████████████████████████
  320 restarts  → batched, priority-ordered,
  rate-limited to protect shared storage

1 site lost
  ████████████████████████████████████████████████████
  → not a restart problem.
  second cluster + explicit failover.


─ capacity check the platform performs ──────────
  lose rack r12 → restart 48 VMs?  yes · 62s
  lose rack r13 → restart 48 VMs?  no · insufficient headroom
  new HA placement that would break the
  reserve is refused at admission.
Topology awareness

Label the metal. Then let the fabric respect it.

Placement that only knows "this host has more free memory" will happily put all three replicas of your database on one rack. Topology labels make location a first-class property of every node.

THE LABEL HIERARCHY

  site
   └── cluster
        └── zone     room · power feed · cooling
             └── rack
                  └── node

  set per node, carried in every heartbeat,
  visible to the console and to agents

EXTRA LABELS — same mechanism
  gpu=h100   nvme=true   storage=local-fast
  used as constraints and selectors by
  placement policy and by agents


UNLABELLED NODES
  default to their own rack → degrades
  gracefully to no topology awareness

What becomes fault-domain aware

  • Spread groups. Members of a service group are placed in different domains, hard or soft.
  • N+1 reserve. Enough free capacity is held across surviving domains to restart everything from the largest one.
  • Prioritised restart. After a rack loss, critical workloads come back first, at a rate the storage can absorb.
  • Quorum placement. Configuration consensus members are spread across domains so losing one does not cost you quorum.
  • Replication targets. Replicated local disks are written to a node in a different domain.
  • Rack-level drain. Take a whole domain out of service with controlled concurrency.
Multi-cluster and multi-site

A cluster is a failure domain. Design your estate as a set of them.

Sites are never part of the same cluster. WAN latency and partitions break storage-lease timing and consensus timing — pretending otherwise is how multi-site hypervisors lose data. Sites are linked instead.

Independent autonomy

Each site runs a full cluster. If the WAN drops, that site is completely unaffected. No split-brain to argue about, because the domains never shared a quorum.

Federated view

A single console across sites, read-only across boundaries. See capacity, inventory and events everywhere without pretending it is one coherent cluster.

Deliberate disaster recovery

Asynchronous disk replication between sites, plus configuration checkpoints. Site failover is an explicit, operator-approved action with a runbook — not something a flapping link triggers at 03:00.

Local media per site

Each site mirrors the image and image library locally, so provisioning and upgrades do not depend on a distant link — and low-cost sites need no WAN bandwidth.

Honest boundary. There is no automatic cross-site high availability, and there will not be. A cluster never restarts another cluster's workloads, because they do not share a fencing domain. Disaster recovery is an operation you perform deliberately.

Operations at any size

The habits that keep a fabric healthy.

Monitor the consensus core

Alert on a shrunken core, replica lag and configuration checkpoint age. Those are the conditions that precede a cold start.

Test your restores

Keep checkpoints off-cluster and restore from them on a schedule. An untested restore path is a guess.

Separate credentials

Distinct cluster tokens for administrative and node traffic. A shared secret means a single leak is a cluster-wide one.

Version-aware upgrades

Ship the message bus with the platform and upgrade it through the same rolling reboot. Finish rollouts promptly — mixed images are expected, not indefinite.

Capacity-plan per site

Nodes, workload density, storage throughput and control-plane bandwidth. The last one is small — keep it that way, and keep migration traffic off the management path.

Run the drills

Node loss, domain loss, consensus member loss, storage latency spikes, cold start. Measure recovery times and record them. Repeat quarterly.

Send us your growth plan.

Node counts today and in three years, where the racks and power feeds are, what your recovery objectives are, and which workloads cannot leave the site. We will model the topology and tell you where the fault domains should be.