Skip to content
User Guide

Multi-Core Single-Region Deployment (HA Cluster)

Three or more mistershell/core containers in one region, sharing external data services and fronted by a load balancer. Any single core can fail without taking the workspace down.

This is one of four supported topologies. For the shared reference (environment variables, TLS, web-session flags, worker registration, scaling, troubleshooting), see the Deployment overview.

Components involved

ComponentWho runs itRole
≥3 core containersYou (MisterShell)Identical cores in the same region. Their embedded message buses self-cluster into a mesh; exactly one core is elected leader and schedules tasks.
Load balancer (LB)You (external)A single virtual hostname in front of the cores, health-checked, that browsers and workers connect to.
External databaseYou (external)Shared relational store — the single source of truth for all cores. Mandatory.
External app-state store (App Redis)You (external)Shared live-session / lock / coordination store. Mandatory, and must run with maxmemory-policy noeviction.
Cache store (Cache Redis)You (external or embedded)Per-core hot cache. Optional to share — each core may keep its own.
Shared object storeYou (external)S3 or Azure bucket for session recordings, so a recording finalized on any core lands in one place.
Remote workersYou (optional)Add to reach resources outside the core’s network zone.

The cores are stateless-ish: all coherence-sensitive state lives in the shared database and App Redis, so a core can be added, removed, or restarted freely.

What is NOT under MisterShell’s responsibility

This is the key boundary of an HA deployment — MisterShell makes the application tier highly available, but it relies on you to make the surrounding infrastructure highly available:

  • The load balancer / GSLB. MisterShell never load-balances itself. You run a health-checked LB and point MISTERSHELL_URL (and your users’ browser URL) at it. Its health check should target GET /health/ready on each core and drop unhealthy cores. Workers reconnect by re-resolving that hostname, so the LB is the failover mechanism for edge connections.
  • Database high availability. The shared database is a single logical dependency — and it is also where cluster leadership is coordinated (a lease a surviving core takes over automatically), so leader failover depends on the database being reachable. Run it as a managed/replicated service with its own failover and backups.
  • Database recovery orchestration. In-product msh restore is intentionally unavailable in every multi-Core topology. Stop all Cores, restore and verify the shared database with PostgreSQL-native tooling, then restart the Cores. See Database Backup and Restore.
  • App Redis high availability. Same — run it replicated, with maxmemory-policy noeviction.
  • Object store durability. The recordings bucket’s replication and lifecycle are the provider’s concern.
  • TLS termination (optional at the LB) and the DNS name + certificate for the shared hostname.
  • Intra-region network reachability between cores on the cluster route port 6222 (TLS-secured mesh).

Architecture

flowchart TB
  edge["Users and Workers"]
  lb["Load Balancer — you run this; health-checks /health/ready"]
  edge --> lb

  subgraph region["region = default · embedded cluster mesh (routes :6222, TLS)"]
    direction LR
    c0["Core 0 · leader"]
    c1["Core 1"]
    c2["Core 2"]
    c0 <--> c1
    c1 <--> c2
    c0 <--> c2
  end

  lb --> c0
  lb --> c1
  lb --> c2

  subgraph shared["Shared external state — you run these, HA is yours"]
    direction LR
    db[("Database")]
    ar[("App Redis · noeviction")]
    os[("Object store")]
  end

  c0 --> shared
  c1 --> shared
  c2 --> shared

Benefits

  • No single point of failure in the application tier. Lose a core and the LB routes around it; a new leader is elected automatically (lease-based election, typically well within ~25 s).
  • Exactly-once scheduling survives failover. Only the leader schedules tasks, and scheduled work is claimed atomically, so a leader handoff never double-runs or drops a tick.
  • Recordings survive node loss. Session recordings are replicated across the region’s cores and finalized to the shared object store — a captured session remains replayable even if the core that hosted it dies mid-session.
  • Rolling upgrades where supported. Rolling-compatible releases replace Cores one at a time while the cluster keeps serving. Check the release notes for cold-upgrade requirements; major-version upgrades and releases requiring a cold upgrade use a coordinated maintenance window. See Upgrades and rollback.
  • A new TLS certificate needs a restart. Certificate and key are read at startup, so after updating them in Settings → System → Advanced Settings, restart every core — rolling, one at a time — for the change to take effect.
  • Horizontal headroom. Add cores to spread API and session load.

Prescriptive deployment guidance

Sizing

  • Run an odd number ≥ 3 of cores. Three tolerates one failure; five tolerates two. Odd counts avoid split quorum in the embedded cluster.
  • Stream replicas default to the region size, so a 3-core region replicates recording/lifecycle streams 3 ways automatically. You do not normally set NATS_STREAM_REPLICAS.

Per-core environment

Every core shares the same DB_ENCRYPTION_KEY, DATABASE_URL, and APP_REDIS_URL, and gets its own identity:

VariablePurposeExample (Core 0)
THIS_CORE_URLThis core’s inter-core address (how peers reach it on the route port). A non-localhost value switches the core into multi-node mode. This is not the address workers use.http://core-0.internal:8000
THIS_CORE_REGIONRegion name = cluster name. Same value on every core in the region.default
NODE_IDStable, unique lease-holder identity. Derived from THIS_CORE_URL if unset; set explicitly on ordered platforms.core-0
DATABASE_URLShared external database (mandatory; a localhost value here refuses to boot in multi-node).postgresql+asyncpg://…@db:5432/mistershell_db
APP_REDIS_URLShared external App Redis (mandatory, noeviction).redis://app-redis:6380/0
DB_ENCRYPTION_KEYSame secret on every core.your-secret-key

EMBEDDED_WORKER is off in multi-node mode and cannot be turned on — a core set to EMBEDDED_WORKER=true with a peer THIS_CORE_URL refuses to start. Run remote workers to process tasks.

Docker Compose (single-host demo of a 3-core region)

Use this to validate the topology on one host; in production the cores belong on separate hosts behind a real LB.

version: '3.8'
x-core: &core
  image: mistershell/core:latest
  environment: &core-env
    DB_ENCRYPTION_KEY: ${DB_ENCRYPTION_KEY}
    THIS_CORE_REGION: default
    DATABASE_URL: postgresql+asyncpg://postgres:postgres@db:5432/mistershell_db
    APP_REDIS_URL: redis://app-redis:6379/0
  depends_on: [db, app-redis]
  restart: unless-stopped

services:
  core-0:
    <<: *core
    environment:
      <<: *core-env
      THIS_CORE_URL: http://core-0:8000
      NODE_ID: core-0
  core-1:
    <<: *core
    environment:
      <<: *core-env
      THIS_CORE_URL: http://core-1:8000
      NODE_ID: core-1
  core-2:
    <<: *core
    environment:
      <<: *core-env
      THIS_CORE_URL: http://core-2:8000
      NODE_ID: core-2

  # Your LB (nginx/HAProxy/cloud LB) fronts core-0..2 on :443. Not shown.
  db:
    image: postgres:17
    environment:
      POSTGRES_DB: mistershell_db
      POSTGRES_PASSWORD: postgres
  app-redis:
    image: redis:7
    command: ["redis-server", "--maxmemory-policy", "noeviction"]

Run the cores as a StatefulSet with parallel pod management (so all pods come up together and form quorum), a headless Service for stable pod DNS (the mesh addresses peers by pod name), a LoadBalancer Service as the shared entry point, and a PodDisruptionBudget so a voluntary drain never drops the cluster below a majority. The manifests below are complete — adjust the image tag, storage size, and your Secret and apply them as-is.

apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: mistershell-core
  labels: { app: mistershell-core }
spec:
  serviceName: mistershell-core     # must match the headless Service below
  replicas: 3
  podManagementPolicy: Parallel     # all pods start together so the cluster forms quorum
  selector:
    matchLabels: { app: mistershell-core }
  template:
    metadata:
      labels: { app: mistershell-core }
    spec:
      containers:
        - name: core
          image: mistershell/core:latest
          ports:
            - { name: https, containerPort: 443 }
            - { name: http, containerPort: 80 }
          env:
            - name: NODE_ID                 # stable, unique per pod
              valueFrom: { fieldRef: { fieldPath: metadata.name } }
            - name: THIS_CORE_URL           # inter-core address = pod DNS
              value: "http://$(NODE_ID).mistershell-core:8000"
            - name: THIS_CORE_REGION
              value: "default"
            - name: DB_ENCRYPTION_KEY
              valueFrom: { secretKeyRef: { name: mistershell, key: db-encryption-key } }
            - name: DATABASE_URL            # external, shared, mandatory
              valueFrom: { secretKeyRef: { name: mistershell, key: database-url } }
            - name: APP_REDIS_URL           # external, shared, noeviction, mandatory
              valueFrom: { secretKeyRef: { name: mistershell, key: app-redis-url } }
          readinessProbe:
            httpGet: { path: /health/ready, port: 443, scheme: HTTPS }
          livenessProbe:
            httpGet: { path: /health/, port: 443, scheme: HTTPS }
          volumeMounts:
            - { name: data, mountPath: /data }
  volumeClaimTemplates:
    - metadata: { name: data }
      spec:
        accessModes: ["ReadWriteOnce"]
        resources: { requests: { storage: 32Gi } }  # the embedded event store alone can grow to 16GB on this volume
---
apiVersion: v1
kind: Service
metadata:
  name: mistershell-core            # headless — stable per-pod DNS for the mesh
spec:
  clusterIP: None
  selector: { app: mistershell-core }
  ports:
    - { name: cluster, port: 6222 }
    # Single-region only. If you reuse this manifest per region in a multi-region
    # topology, also publish the cross-region gateway port:
    #   - { name: gateway, port: 7222 }
---
apiVersion: v1
kind: Service
metadata:
  name: mistershell                 # the shared entry point (your LB / MISTERSHELL_URL)
spec:
  type: LoadBalancer
  selector: { app: mistershell-core }
  ports:
    - { name: https, port: 443, targetPort: 443 }
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: mistershell-core
spec:
  minAvailable: 2                    # ⌊N/2⌋+1 (2 of 3) — keep a majority during drains
  selector:
    matchLabels: { app: mistershell-core }

Create the referenced Secret (mistershell) with db-encryption-key, database-url, and app-redis-url out of band. If you host web-application sessions, add the headless-Chrome security context to the pod.

Recordings

Configure a shared cloud recording store (S3 or Azure) under Govern → Recording Policy → Stores and point your recording rules at it, so recordings finalized on any core land in one bucket. The built-in Local store writes to each core’s own disk and is not appropriate for a cluster. See Recording Policy → Add a cloud recording store.

Configure the load balancer

  • Health check: GET /health/ready on :443 (or :80), per core; remove a core from rotation when it fails.
  • Allow WebSocket upgrades and set a generous idle timeout (5+ minutes) for long-lived sessions.
  • Sticky sessions are not required — any core can serve any request.
  • Point MISTERSHELL_URL (used by workers) and your users’ browser URL at the LB hostname.

When to choose something else

If you need…Move to
A presence in several regionsDistributed (one core each) or Distributed with HA (redundant per region)
Just a quick single-host workspaceAll in One