Skip to content
User Guide

Multi-Core, Multi-Region (Distributed with HA)

The full topology: three or more cores per region, across two or more regions, federated into one cross-region mesh over a shared central data tier. It is one HA fabric that spans regions — a single elected leader and one shared central data tier coordinate the whole estate — with each region internally redundant and geographically local to its users.

This is one of four supported topologies. For the shared reference (environment variables, TLS, web-session flags, worker registration, scaling, troubleshooting), see the Deployment overview. This page combines the HA Cluster and Distributed models — read those first for the details this page builds on.

Components involved

ComponentWho runs itRole
≥3 cores per regionYou (MisterShell)Each region’s cores self-cluster (routes on :6222); regions federate through gateways (:7222) into one super-cluster. The whole fabric elects a single leader and shares one central data tier — it is one HA fabric spanning regions, not independent clusters.
Per-region load balancerYou (external)A regional VIP in front of that region’s cores, health-checked.
GSLBYou (external)Global steering across the regional VIPs, with health-based failover.
Central databaseYou (external)One shared relational store. Mandatory.
Central app-state store (App Redis)You (external)One shared coordination store. Mandatory, noeviction.
Central object storeYou (external)One bucket for recordings from every region.
Per-region workersYou (optional)Attach to their regional cores.

What is NOT under MisterShell’s responsibility

This topology has the widest infrastructure boundary — everything from both simpler topologies applies:

  • The GSLB and the per-region load balancers. MisterShell load-balances neither across regions nor within one. You run the global steering and each regional VIP; MISTERSHELL_URL resolves through them and workers re-resolve on reconnect, so your LB/GSLB layer is the failover mechanism at both scopes.
  • Database and App Redis high availability. Still a single shared, central tier — its replication, failover, backups, and cross-region latency are yours. This is the one component that is not regionally redundant, and the central database is also where fabric-wide leadership is coordinated, so treat its HA as the foundation of the whole deployment (managed multi-AZ / globally-distributed).
  • Object store durability.
  • Network reachability: the cluster route port 6222 between cores within each region, and the gateway port 7222 between regions (both TLS).
  • TLS termination, DNS, and certificates for every regional entry point and the global name.

Architecture

flowchart TB
  users["Users / Workers"]
  gslb["GSLB · global steering + health failover"]
  users --> gslb
  gslb --> lbUS["LB · us-east"]
  gslb --> lbEU["LB · eu-west"]
  subgraph rUS["region us-east · routes :6222 · R = 3"]
    direction LR
    u0["Core 0"]
    u1["Core 1"]
    u2["Core 2"]
    u0 <--> u1
    u1 <--> u2
    u0 <--> u2
  end
  subgraph rEU["region eu-west · routes :6222 · R = 3"]
    direction LR
    e0["Core 0"]
    e1["Core 1"]
    e2["Core 2"]
    e0 <--> e1
    e1 <--> e2
    e0 <--> e2
  end
  lbUS --> u0
  lbUS --> u1
  lbUS --> u2
  lbEU --> e0
  lbEU --> e1
  lbEU --> e2
  rUS <-. "gateway :7222 · super-cluster federation" .-> rEU
  state[("Central external state<br/>Database · App Redis · Object store<br/>coordinates ONE fabric-wide leader · HA is yours")]
  u1 --> state
  e1 --> state

Benefits

  • HA within every region. Lose a core and that region’s LB routes around it; the region’s remaining cores keep serving. If the lost core held the fabric’s leadership, a new leader is elected from anywhere in the fabric — the same failover an HA Cluster provides, across regions.
  • Geographic distribution. Users and workers hit their nearest region for low-latency sessions and local task execution.
  • Per-region recording durability. Each region replicates its recording streams R = region size, so a captured session survives the loss of a core in its own region, then finalizes to the central object store.
  • Rolling upgrades that bound blast radius. Compatible releases roll across the fabric one core at a time. Workers connected to the cycled Core re-register through the regional load balancer; their ongoing tasks and sessions close on disconnect. Other regions keep serving their users. Check the release notes for cold-upgrade requirements and follow Upgrades and rollback.
  • Graceful degradation. A whole region going dark is contained by the GSLB; surviving regions keep serving as long as the central data tier is reachable.

Prescriptive deployment guidance

Sizing and quorum

  • Odd number ≥ 3 cores per region. Three cores per region let a region lose one of its own cores and keep serving on the other two, and replicate that region’s recording streams R = region size. Fabric-wide leadership is coordinated through the shared central database (one leader for the whole fabric), so it survives core loss anywhere as long as the central tier is reachable. Keep each region’s core count odd for the embedded event store, whose replicated groups elect by majority.
  • Stream replicas default to each region’s size; leave NATS_STREAM_REPLICAS unset so a 3-core region replicates 3 ways and a 5-core region replicates 5 ways automatically.

Per-core environment

Combine the two axes: region identity (shared within a region) and node identity (unique per core). Every core across the whole estate shares DB_ENCRYPTION_KEY, the central DATABASE_URL, and the central APP_REDIS_URL.

Variableus-east / Core 0us-east / Core 1eu-west / Core 0
THIS_CORE_URLhttp://use1-c0.internal:8000http://use1-c1.internal:8000http://euw1-c0.internal:8000
THIS_CORE_REGIONus-eastus-easteu-west
NODE_IDuse1-core-0use1-core-1euw1-core-0
DATABASE_URLcentral (same everywhere)centralcentral
APP_REDIS_URLcentral (same everywhere)centralcentral

NODE_ID must be globally unique and stable; THIS_CORE_REGION groups cores into their regional mesh.

Run one StatefulSet per region — typically one cluster (or namespace) per region so each region’s pods form their own intra-region routing mesh, federated into the one fabric by gateways:

  • Each StatefulSet: parallel pod management, a headless Service (intra-region mesh DNS), a regional LoadBalancer Service (that region’s VIP), and a PodDisruptionBudget with minAvailable = ⌊N/2⌋+1 (2 of 3), so a node drain never drops a region below its serving floor or its event store’s replication majority.
  • Set THIS_CORE_REGION to the region, NODE_ID from the pod name (prefixed per region for global uniqueness), and THIS_CORE_URL from the stable pod DNS.
  • Point every region’s DATABASE_URL / APP_REDIS_URL at the same central managed services.

Use the single-region StatefulSet + Services + PodDisruptionBudget from the HA Cluster page as your per-region template; deploy one copy per region, changing THIS_CORE_REGION and prefixing NODE_ID per region so identities stay globally unique.

Network

  • Within each region: cluster route port 6222 between that region’s cores.
  • Between regions: cross-region gateway port 7222 between the regional core subnets.
  • Every region reaches the central database, App Redis, and object store.

GSLB, LBs, and recordings

  • Front each region with its own health-checked LB (GET /health/ready, WebSocket upgrades allowed, generous idle timeout), and put a GSLB over the regional VIPs for global steering + failover.
  • Point MISTERSHELL_URL at the GSLB name.
  • Configure a shared cloud recording store (S3/Azure) under Govern → Recording Policy → Stores and point your recording rules at it, so recordings from every region land in one place. See Recording Policy → Add a cloud recording store.

Upgrades

It is one fabric, so an upgrade is one operation over the whole fabric.

Rolling-compatible releases let you replace cores one at a time; each new core reconciles shared state on start and rejoins the mesh, so the fabric keeps serving throughout. You can sequence the rollout region by region — cycle one region’s cores while the other regions keep serving their users unaffected — to bound the blast radius. This is one rolling upgrade of one fabric, not independent per-region upgrades.

Cold upgrades, including major-version upgrades and releases marked as requiring one in the release notes, replace the whole fabric in one planned window. Quiesce and drain, stop every Core across every region, then bring up the target version together. Follow Upgrades and rollback for backup and recovery requirements.

When to choose something else

If you need…Move to
Multiple regions but no per-region redundancyDistributed
Redundancy in a single regionHA Cluster
A single-host workspaceAll in One