Multi-Core Single-Region Deployment (HA Cluster)
Three or more mistershell/core containers in one region, sharing external data services and fronted by a load balancer. Any single core can fail without taking the workspace down.
This is one of four supported topologies. For the shared reference (environment variables, TLS, web-session flags, worker registration, scaling, troubleshooting), see the Deployment overview.
Components involved
| Component | Who runs it | Role |
|---|---|---|
| ≥3 core containers | You (MisterShell) | Identical cores in the same region. Their embedded message buses self-cluster into a mesh; exactly one core is elected leader and schedules tasks. |
| Load balancer (LB) | You (external) | A single virtual hostname in front of the cores, health-checked, that browsers and workers connect to. |
| External database | You (external) | Shared relational store — the single source of truth for all cores. Mandatory. |
| External app-state store (App Redis) | You (external) | Shared live-session / lock / coordination store. Mandatory, and must run with maxmemory-policy noeviction. |
| Cache store (Cache Redis) | You (external or embedded) | Per-core hot cache. Optional to share — each core may keep its own. |
| Shared object store | You (external) | S3 or Azure bucket for session recordings, so a recording finalized on any core lands in one place. |
| Remote workers | You (optional) | Add to reach resources outside the core’s network zone. |
The cores are stateless-ish: all coherence-sensitive state lives in the shared database and App Redis, so a core can be added, removed, or restarted freely.
What is NOT under MisterShell’s responsibility
This is the key boundary of an HA deployment — MisterShell makes the application tier highly available, but it relies on you to make the surrounding infrastructure highly available:
- The load balancer / GSLB. MisterShell never load-balances itself. You run a health-checked LB and point
MISTERSHELL_URL(and your users’ browser URL) at it. Its health check should targetGET /health/readyon each core and drop unhealthy cores. Workers reconnect by re-resolving that hostname, so the LB is the failover mechanism for edge connections. - Database high availability. The shared database is a single logical dependency — and it is also where cluster leadership is coordinated (a lease a surviving core takes over automatically), so leader failover depends on the database being reachable. Run it as a managed/replicated service with its own failover and backups.
- Database recovery orchestration. In-product
msh restoreis intentionally unavailable in every multi-Core topology. Stop all Cores, restore and verify the shared database with PostgreSQL-native tooling, then restart the Cores. See Database Backup and Restore. - App Redis high availability. Same — run it replicated, with
maxmemory-policy noeviction. - Object store durability. The recordings bucket’s replication and lifecycle are the provider’s concern.
- TLS termination (optional at the LB) and the DNS name + certificate for the shared hostname.
- Intra-region network reachability between cores on the cluster route port
6222(TLS-secured mesh).
Architecture
flowchart TB
edge["Users and Workers"]
lb["Load Balancer — you run this; health-checks /health/ready"]
edge --> lb
subgraph region["region = default · embedded cluster mesh (routes :6222, TLS)"]
direction LR
c0["Core 0 · leader"]
c1["Core 1"]
c2["Core 2"]
c0 <--> c1
c1 <--> c2
c0 <--> c2
end
lb --> c0
lb --> c1
lb --> c2
subgraph shared["Shared external state — you run these, HA is yours"]
direction LR
db[("Database")]
ar[("App Redis · noeviction")]
os[("Object store")]
end
c0 --> shared
c1 --> shared
c2 --> shared
Benefits
- No single point of failure in the application tier. Lose a core and the LB routes around it; a new leader is elected automatically (lease-based election, typically well within ~25 s).
- Exactly-once scheduling survives failover. Only the leader schedules tasks, and scheduled work is claimed atomically, so a leader handoff never double-runs or drops a tick.
- Recordings survive node loss. Session recordings are replicated across the region’s cores and finalized to the shared object store — a captured session remains replayable even if the core that hosted it dies mid-session.
- Rolling upgrades where supported. Rolling-compatible releases replace Cores one at a time while the cluster keeps serving. Check the release notes for cold-upgrade requirements; major-version upgrades and releases requiring a cold upgrade use a coordinated maintenance window. See Upgrades and rollback.
- A new TLS certificate needs a restart. Certificate and key are read at startup, so after updating them in Settings → System → Advanced Settings, restart every core — rolling, one at a time — for the change to take effect.
- Horizontal headroom. Add cores to spread API and session load.
Prescriptive deployment guidance
Sizing
- Run an odd number ≥ 3 of cores. Three tolerates one failure; five tolerates two. Odd counts avoid split quorum in the embedded cluster.
- Stream replicas default to the region size, so a 3-core region replicates recording/lifecycle streams 3 ways automatically. You do not normally set
NATS_STREAM_REPLICAS.
Per-core environment
Every core shares the same DB_ENCRYPTION_KEY, DATABASE_URL, and APP_REDIS_URL, and gets its own identity:
| Variable | Purpose | Example (Core 0) |
|---|---|---|
THIS_CORE_URL | This core’s inter-core address (how peers reach it on the route port). A non-localhost value switches the core into multi-node mode. This is not the address workers use. | http://core-0.internal:8000 |
THIS_CORE_REGION | Region name = cluster name. Same value on every core in the region. | default |
NODE_ID | Stable, unique lease-holder identity. Derived from THIS_CORE_URL if unset; set explicitly on ordered platforms. | core-0 |
DATABASE_URL | Shared external database (mandatory; a localhost value here refuses to boot in multi-node). | postgresql+asyncpg://…@db:5432/mistershell_db |
APP_REDIS_URL | Shared external App Redis (mandatory, noeviction). | redis://app-redis:6380/0 |
DB_ENCRYPTION_KEY | Same secret on every core. | your-secret-key |
EMBEDDED_WORKERis off in multi-node mode and cannot be turned on — a core set toEMBEDDED_WORKER=truewith a peerTHIS_CORE_URLrefuses to start. Run remote workers to process tasks.
Docker Compose (single-host demo of a 3-core region)
Use this to validate the topology on one host; in production the cores belong on separate hosts behind a real LB.
version: '3.8'
x-core: &core
image: mistershell/core:latest
environment: &core-env
DB_ENCRYPTION_KEY: ${DB_ENCRYPTION_KEY}
THIS_CORE_REGION: default
DATABASE_URL: postgresql+asyncpg://postgres:postgres@db:5432/mistershell_db
APP_REDIS_URL: redis://app-redis:6379/0
depends_on: [db, app-redis]
restart: unless-stopped
services:
core-0:
<<: *core
environment:
<<: *core-env
THIS_CORE_URL: http://core-0:8000
NODE_ID: core-0
core-1:
<<: *core
environment:
<<: *core-env
THIS_CORE_URL: http://core-1:8000
NODE_ID: core-1
core-2:
<<: *core
environment:
<<: *core-env
THIS_CORE_URL: http://core-2:8000
NODE_ID: core-2
# Your LB (nginx/HAProxy/cloud LB) fronts core-0..2 on :443. Not shown.
db:
image: postgres:17
environment:
POSTGRES_DB: mistershell_db
POSTGRES_PASSWORD: postgres
app-redis:
image: redis:7
command: ["redis-server", "--maxmemory-policy", "noeviction"]
Kubernetes (recommended for production)
Run the cores as a StatefulSet with parallel pod management (so all pods come up together and form quorum), a headless Service for stable pod DNS (the mesh addresses peers by pod name), a LoadBalancer Service as the shared entry point, and a PodDisruptionBudget so a voluntary drain never drops the cluster below a majority. The manifests below are complete — adjust the image tag, storage size, and your Secret and apply them as-is.
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: mistershell-core
labels: { app: mistershell-core }
spec:
serviceName: mistershell-core # must match the headless Service below
replicas: 3
podManagementPolicy: Parallel # all pods start together so the cluster forms quorum
selector:
matchLabels: { app: mistershell-core }
template:
metadata:
labels: { app: mistershell-core }
spec:
containers:
- name: core
image: mistershell/core:latest
ports:
- { name: https, containerPort: 443 }
- { name: http, containerPort: 80 }
env:
- name: NODE_ID # stable, unique per pod
valueFrom: { fieldRef: { fieldPath: metadata.name } }
- name: THIS_CORE_URL # inter-core address = pod DNS
value: "http://$(NODE_ID).mistershell-core:8000"
- name: THIS_CORE_REGION
value: "default"
- name: DB_ENCRYPTION_KEY
valueFrom: { secretKeyRef: { name: mistershell, key: db-encryption-key } }
- name: DATABASE_URL # external, shared, mandatory
valueFrom: { secretKeyRef: { name: mistershell, key: database-url } }
- name: APP_REDIS_URL # external, shared, noeviction, mandatory
valueFrom: { secretKeyRef: { name: mistershell, key: app-redis-url } }
readinessProbe:
httpGet: { path: /health/ready, port: 443, scheme: HTTPS }
livenessProbe:
httpGet: { path: /health/, port: 443, scheme: HTTPS }
volumeMounts:
- { name: data, mountPath: /data }
volumeClaimTemplates:
- metadata: { name: data }
spec:
accessModes: ["ReadWriteOnce"]
resources: { requests: { storage: 32Gi } } # the embedded event store alone can grow to 16GB on this volume
---
apiVersion: v1
kind: Service
metadata:
name: mistershell-core # headless — stable per-pod DNS for the mesh
spec:
clusterIP: None
selector: { app: mistershell-core }
ports:
- { name: cluster, port: 6222 }
# Single-region only. If you reuse this manifest per region in a multi-region
# topology, also publish the cross-region gateway port:
# - { name: gateway, port: 7222 }
---
apiVersion: v1
kind: Service
metadata:
name: mistershell # the shared entry point (your LB / MISTERSHELL_URL)
spec:
type: LoadBalancer
selector: { app: mistershell-core }
ports:
- { name: https, port: 443, targetPort: 443 }
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: mistershell-core
spec:
minAvailable: 2 # ⌊N/2⌋+1 (2 of 3) — keep a majority during drains
selector:
matchLabels: { app: mistershell-core }
Create the referenced Secret (mistershell) with db-encryption-key, database-url, and app-redis-url out of band. If you host web-application sessions, add the headless-Chrome security context to the pod.
Recordings
Configure a shared cloud recording store (S3 or Azure) under Govern → Recording Policy → Stores and point your recording rules at it, so recordings finalized on any core land in one bucket. The built-in Local store writes to each core’s own disk and is not appropriate for a cluster. See Recording Policy → Add a cloud recording store.
Configure the load balancer
- Health check:
GET /health/readyon:443(or:80), per core; remove a core from rotation when it fails. - Allow WebSocket upgrades and set a generous idle timeout (5+ minutes) for long-lived sessions.
- Sticky sessions are not required — any core can serve any request.
- Point
MISTERSHELL_URL(used by workers) and your users’ browser URL at the LB hostname.
When to choose something else
| If you need… | Move to |
|---|---|
| A presence in several regions | Distributed (one core each) or Distributed with HA (redundant per region) |
| Just a quick single-host workspace | All in One |