Multi-Core, Multi-Region (Distributed with HA)
The full topology: three or more cores per region, across two or more regions, federated into one cross-region mesh over a shared central data tier. It is one HA fabric that spans regions — a single elected leader and one shared central data tier coordinate the whole estate — with each region internally redundant and geographically local to its users.
This is one of four supported topologies. For the shared reference (environment variables, TLS, web-session flags, worker registration, scaling, troubleshooting), see the Deployment overview. This page combines the HA Cluster and Distributed models — read those first for the details this page builds on.
Components involved
| Component | Who runs it | Role |
|---|---|---|
| ≥3 cores per region | You (MisterShell) | Each region’s cores self-cluster (routes on :6222); regions federate through gateways (:7222) into one super-cluster. The whole fabric elects a single leader and shares one central data tier — it is one HA fabric spanning regions, not independent clusters. |
| Per-region load balancer | You (external) | A regional VIP in front of that region’s cores, health-checked. |
| GSLB | You (external) | Global steering across the regional VIPs, with health-based failover. |
| Central database | You (external) | One shared relational store. Mandatory. |
| Central app-state store (App Redis) | You (external) | One shared coordination store. Mandatory, noeviction. |
| Central object store | You (external) | One bucket for recordings from every region. |
| Per-region workers | You (optional) | Attach to their regional cores. |
What is NOT under MisterShell’s responsibility
This topology has the widest infrastructure boundary — everything from both simpler topologies applies:
- The GSLB and the per-region load balancers. MisterShell load-balances neither across regions nor within one. You run the global steering and each regional VIP;
MISTERSHELL_URLresolves through them and workers re-resolve on reconnect, so your LB/GSLB layer is the failover mechanism at both scopes. - Database and App Redis high availability. Still a single shared, central tier — its replication, failover, backups, and cross-region latency are yours. This is the one component that is not regionally redundant, and the central database is also where fabric-wide leadership is coordinated, so treat its HA as the foundation of the whole deployment (managed multi-AZ / globally-distributed).
- Object store durability.
- Network reachability: the cluster route port
6222between cores within each region, and the gateway port7222between regions (both TLS). - TLS termination, DNS, and certificates for every regional entry point and the global name.
Architecture
flowchart TB
users["Users / Workers"]
gslb["GSLB · global steering + health failover"]
users --> gslb
gslb --> lbUS["LB · us-east"]
gslb --> lbEU["LB · eu-west"]
subgraph rUS["region us-east · routes :6222 · R = 3"]
direction LR
u0["Core 0"]
u1["Core 1"]
u2["Core 2"]
u0 <--> u1
u1 <--> u2
u0 <--> u2
end
subgraph rEU["region eu-west · routes :6222 · R = 3"]
direction LR
e0["Core 0"]
e1["Core 1"]
e2["Core 2"]
e0 <--> e1
e1 <--> e2
e0 <--> e2
end
lbUS --> u0
lbUS --> u1
lbUS --> u2
lbEU --> e0
lbEU --> e1
lbEU --> e2
rUS <-. "gateway :7222 · super-cluster federation" .-> rEU
state[("Central external state<br/>Database · App Redis · Object store<br/>coordinates ONE fabric-wide leader · HA is yours")]
u1 --> state
e1 --> state
Benefits
- HA within every region. Lose a core and that region’s LB routes around it; the region’s remaining cores keep serving. If the lost core held the fabric’s leadership, a new leader is elected from anywhere in the fabric — the same failover an HA Cluster provides, across regions.
- Geographic distribution. Users and workers hit their nearest region for low-latency sessions and local task execution.
- Per-region recording durability. Each region replicates its recording streams
R = region size, so a captured session survives the loss of a core in its own region, then finalizes to the central object store. - Rolling upgrades that bound blast radius. Compatible releases roll across the fabric one core at a time. Workers connected to the cycled Core re-register through the regional load balancer; their ongoing tasks and sessions close on disconnect. Other regions keep serving their users. Check the release notes for cold-upgrade requirements and follow Upgrades and rollback.
- Graceful degradation. A whole region going dark is contained by the GSLB; surviving regions keep serving as long as the central data tier is reachable.
Prescriptive deployment guidance
Sizing and quorum
- Odd number ≥ 3 cores per region. Three cores per region let a region lose one of its own cores and keep serving on the other two, and replicate that region’s recording streams
R = region size. Fabric-wide leadership is coordinated through the shared central database (one leader for the whole fabric), so it survives core loss anywhere as long as the central tier is reachable. Keep each region’s core count odd for the embedded event store, whose replicated groups elect by majority. - Stream replicas default to each region’s size; leave
NATS_STREAM_REPLICASunset so a 3-core region replicates 3 ways and a 5-core region replicates 5 ways automatically.
Per-core environment
Combine the two axes: region identity (shared within a region) and node identity (unique per core). Every core across the whole estate shares DB_ENCRYPTION_KEY, the central DATABASE_URL, and the central APP_REDIS_URL.
| Variable | us-east / Core 0 | us-east / Core 1 | eu-west / Core 0 |
|---|---|---|---|
THIS_CORE_URL | http://use1-c0.internal:8000 | http://use1-c1.internal:8000 | http://euw1-c0.internal:8000 |
THIS_CORE_REGION | us-east | us-east | eu-west |
NODE_ID | use1-core-0 | use1-core-1 | euw1-core-0 |
DATABASE_URL | central (same everywhere) | central | central |
APP_REDIS_URL | central (same everywhere) | central | central |
NODE_ID must be globally unique and stable; THIS_CORE_REGION groups cores into their regional mesh.
Kubernetes (recommended)
Run one StatefulSet per region — typically one cluster (or namespace) per region so each region’s pods form their own intra-region routing mesh, federated into the one fabric by gateways:
- Each StatefulSet: parallel pod management, a headless Service (intra-region mesh DNS), a regional LoadBalancer Service (that region’s VIP), and a PodDisruptionBudget with
minAvailable = ⌊N/2⌋+1(2 of 3), so a node drain never drops a region below its serving floor or its event store’s replication majority. - Set
THIS_CORE_REGIONto the region,NODE_IDfrom the pod name (prefixed per region for global uniqueness), andTHIS_CORE_URLfrom the stable pod DNS. - Point every region’s
DATABASE_URL/APP_REDIS_URLat the same central managed services.
Use the single-region StatefulSet + Services + PodDisruptionBudget from the HA Cluster page as your per-region template; deploy one copy per region, changing THIS_CORE_REGION and prefixing NODE_ID per region so identities stay globally unique.
Network
- Within each region: cluster route port
6222between that region’s cores. - Between regions: cross-region gateway port
7222between the regional core subnets. - Every region reaches the central database, App Redis, and object store.
GSLB, LBs, and recordings
- Front each region with its own health-checked LB (
GET /health/ready, WebSocket upgrades allowed, generous idle timeout), and put a GSLB over the regional VIPs for global steering + failover. - Point
MISTERSHELL_URLat the GSLB name. - Configure a shared cloud recording store (S3/Azure) under Govern → Recording Policy → Stores and point your recording rules at it, so recordings from every region land in one place. See Recording Policy → Add a cloud recording store.
Upgrades
It is one fabric, so an upgrade is one operation over the whole fabric.
Rolling-compatible releases let you replace cores one at a time; each new core reconciles shared state on start and rejoins the mesh, so the fabric keeps serving throughout. You can sequence the rollout region by region — cycle one region’s cores while the other regions keep serving their users unaffected — to bound the blast radius. This is one rolling upgrade of one fabric, not independent per-region upgrades.
Cold upgrades, including major-version upgrades and releases marked as requiring one in the release notes, replace the whole fabric in one planned window. Quiesce and drain, stop every Core across every region, then bring up the target version together. Follow Upgrades and rollback for backup and recovery requirements.
When to choose something else
| If you need… | Move to |
|---|---|
| Multiple regions but no per-region redundancy | Distributed |
| Redundancy in a single region | HA Cluster |
| A single-host workspace | All in One |