Skip to content
User Guide

Deployment

MisterShell is distributed as four published container images. You pull them directly from the registry.

ImagePurpose
mistershell/core:latestThe control plane — web UI, API, event bus, session gateway, and an embedded worker. This is what your users connect to.
mistershell/worker:latestA remote worker for distributed deployments. Install one per site, region, or network segment.
mistershell/proxy:latestA public-facing session proxy that lets external guests reach a shared session without exposing your core. Optional; requires the Session Proxy add-on — see Proxies.
mistershell/sensor:latestA remote intrusion-detection probe that passively watches network traffic and raises alerts. Optional; requires the IDS Sensors add-on — see Sensors.

Every deployment uses core (and, for distributed topologies, worker). The proxy and sensor are optional add-ons you install only if you use external session sharing or intrusion detection, and each needs its own license add-on.

Pin every image to the chosen release tag (mistershell/core:<version>, for example) in production so that upgrades are deliberate.

This page is the shared reference for every deployment: prerequisites, environment variables, TLS, web sessions, worker registration, scaling, and troubleshooting. Pick your topology below — each topology page has the components, boundaries, diagram, benefits, and step-by-step guidance specific to it, and links back here for the shared bits.

Choose a deployment topology

MisterShell supports four topologies. They differ along two axes: how many cores per region (redundancy) and how many regions (geographic reach).

TopologyCoresRegionsIn-region HAGeographicExternal services you run
All in One11——None (all embedded)
HA Cluster≥31✅—Database, App Redis, LB, object store
Distributed1 / region≥2—✅Central database, App Redis, GSLB, object store
Distributed with HA≥3 / region≥2✅✅Central database, App Redis, GSLB + per-region LBs, object store
  • All in One — one container, everything embedded. Evaluation, proofs-of-concept, small single-zone estates.
  • HA Cluster — multiple cores in one region, no single point of failure in the application tier.
  • Distributed — one core per region for geographic locality, without per-region redundancy.
  • Distributed with HA — redundant cores in each of several regions; the full topology.

A common boundary applies to every multi-core topology: MisterShell makes its own application tier highly available (leader election, mesh, replicated recordings), but it depends on you to run the load balancer / GSLB, the database, the app-state store (App Redis), and the object store — and to make those highly available. Note that cluster leadership itself is coordinated through the database (a lease that a surviving core takes over automatically), so the database’s availability is what keeps failover working. Each topology page spells out exactly what falls on which side of that line.

Upgrades and rollback

Back up the MisterShell database and verify that the backup is readable before every upgrade. Also preserve DB_ENCRYPTION_KEY and any externally managed recording store according to your retention policy. A successful application startup or schema migration is not a substitute for a restorable backup. See Database Backup and Restore for the scheduled task, MSH commands, and the different single-Core and multi-Core procedures.

For a single-Core deployment, back up the data, replace the container with the new image, and start it. Pending database migrations and registry synchronization run automatically at startup.

Cold and rolling upgrade considerations apply only to multi-Core deployments. Within one major version, eligible clusters support rolling upgrades unless the release notes require a coordinated cold upgrade. For a cold upgrade, stop new sessions, let active sessions finish, stop every old Core and Worker, take the verified backup, and then start only the new version. Multi-Core upgrades across major versions require this coordinated procedure; old and new major versions must not run concurrently.

Rollback after a major-version upgrade means restoring the pre-upgrade database backup and the matching previous-version containers together. Do not use schema downgrades as a data-recovery procedure. Confirm database, recording-store, and encryption-key recovery in a non-production environment before relying on the procedure.

Prerequisites

  • A Linux host with a container runtime (Docker, Podman, or equivalent) per core and per remote worker.
  • DNS / TLS planning: the core server answers on :443 (HTTPS + WebSocket upgrades) and :80. By default :80 serves the full application over plain HTTP (for use behind your own TLS-terminating reverse proxy); set ENABLE_HTTP_REDIRECT=true to make :80 serve only health checks and HTTP → HTTPS redirects. Do not expose :80 to untrusted networks unless the redirect is enabled.
  • If you enable the syslog collector on the core’s embedded worker, also map the collector listener port (514/udp and 514/tcp by default).
  • Native SSH access (ssh/sftp/scp against sessions and the file catalog) listens on the sshgw_port setting (default 2222); map it 1:1 at the host, in Compose, and on the Kubernetes load balancer if you want operators to use it. The gateway is started and supervised by the Core itself — no separate container or process to run.
  • For multi-site deployments, outbound HTTPS/WebSocket connectivity from each worker host back to the core’s public URL.
  • For any multi-core topology, an externally managed relational database and app-state store (App Redis) reachable from every core, plus a load balancer / GSLB in front of the cores. See the topology pages for the exact requirements.

Web sessions (headless Chrome)

Web-application sessions open the target site in a headless Chrome running next to the worker. Chrome is already bundled in both the core and worker images (including the embedded worker inside the core image), so there is nothing to install. Chrome keeps its renderer sandbox on — the isolation boundary between the visited page and the worker — and that sandbox needs to create namespaces, which the container blocks by default.

Requirement — run the core container (and any remote worker that hosts web sessions) with these two settings:

--cap-add SYS_ADMIN --security-opt apparmor=unconfined
  • --cap-add SYS_ADMIN lets the container create the namespaces Chrome’s sandbox needs, while keeping the default seccomp filter on.
  • --security-opt apparmor=unconfined clears Ubuntu 22.04/24.04’s AppArmor user-namespace restriction; it’s a harmless no-op on hosts without AppArmor, so the same two settings work everywhere.

Without them, web sessions fail to start and the logs show Failed to move to new namespace: Operation not permitted. Other session types (SSH, RDP, VNC, cloud, DB) are unaffected. Don’t disable the sandbox with Chrome --no-sandbox — that removes the renderer isolation and is a security regression.

Kubernetes — per core/worker pod securityContext:

securityContext:
  capabilities:
    add: ["SYS_ADMIN"]
# AppArmor: set the container annotation to unconfined, e.g.
#   container.apparmor.security.beta.kubernetes.io/<container>: unconfined

Portainer — deploy the core/worker as a Stack using the Compose YAML in Docker Compose examples; the cap_add / security_opt keys apply as-is. Edit and redeploy the stack to recreate the container with the options (changing them on a running container has no effect). On a Portainer Swarm environment security_opt is ignored, so rely on cap_add: SYS_ADMIN.

Notes:

  • A simpler but broader alternative is --security-opt seccomp=unconfined --security-opt apparmor=unconfined; it works too, but turns off the whole seccomp filter — prefer the SYS_ADMIN form, which keeps it on.
  • Locked-down platforms (Kubernetes restricted/baseline PodSecurity, OpenShift, Docker Swarm, rootless/gVisor) may forbid these options; there, web sessions can’t run and every other feature still works.

Environment variables

Core container — required

VariableDescription
DB_ENCRYPTION_KEYSecret used to encrypt sensitive values stored in the database. Treat this as a credential and back it up. In a multi-core deployment, every core uses the same value.

Core container — optional

VariableDefaultDescription
DATABASE_URLembedded databaseConnection string for the relational database. A localhost / 127.0.0.1 host activates the embedded database; an external host routes to your own service. Mandatory (external) in any multi-core topology.
DB_POOL_SIZE20Persistent database connections. Scale alongside data_collection_concurrency — see Scaling.
DB_MAX_OVERFLOW40Burst database connections under load. Total pool capacity = DB_POOL_SIZE + DB_MAX_OVERFLOW.
CACHE_REDIS_URLembedded cacheConnection string for the high-performance in-memory hot cache. Same localhost rule as DATABASE_URL. May stay per-core even in a cluster.
APP_REDIS_URLembedded app storeConnection string for the app-state store (live sessions, locks, cluster coordination state). Same localhost rule as DATABASE_URL. Mandatory (external, shared) in any multi-core topology, and the external instance must use maxmemory-policy noeviction.
CORS_ORIGINS(empty)Comma-separated list of browser origins allowed for cross-origin API requests. Set this only when the UI and API are served from different origins.
LOG_LEVELINFOOne of DEBUG, INFO, WARNING, ERROR, CRITICAL.
TLS_CERT_CNlocalhostCommon name used for the self-signed bootstrap certificate generated on first start.
ENABLE_HTTP_REDIRECTfalseWhen true, :80 serves only health checks and redirects everything else to HTTPS. When false (the default), :80 serves the full application over plain HTTP — intended for running behind your own TLS-terminating reverse proxy.
EMBEDDED_WORKERtrue (single-node)Start the embedded worker inside the core container. In multi-node mode it defaults off and cannot be turned on — setting it to true there is a startup error; run remote workers instead.

Core container — clustering (multi-core topologies only)

These select and shape a multi-core deployment. On a single-core (All in One) deployment, leave them at their defaults.

VariableDefaultDescription
THIS_CORE_URLhttp://localhost:8000This core’s inter-core address — how peer cores reach it on the cluster route port (:6222) within a region and the gateway port (:7222) across regions. A non-localhost value is what switches the core into multi-node mode. This is not the address workers use — that is MISTERSHELL_URL, which points at your load balancer.
THIS_CORE_REGIONdefaultRegion name; also the cluster name. Same on every core in a region. Must match ^[a-z0-9-]+$ (≤32 chars).
NODE_IDderived from THIS_CORE_URLStable, unique lease-holder identity (used for leader election and the cluster mesh). If unset, derived from the THIS_CORE_URL host and port (e.g. http://core-0.internal:8000 → core-0.internal_8000, unsafe characters mapped to _); set explicitly on ordered platforms (e.g. from a StatefulSet pod name). Must match ^[A-Za-z0-9._-]+$ (1–64 chars).
NATS_STREAM_REPLICASregion sizeReplica count for the region’s event/recording streams (1–9). Defaults to the number of cores in the region, so you normally leave it unset.

Remote worker — required

VariableDescription
MISTERSHELL_URLPublic URL of your core deployment — your load balancer / GSLB hostname (for example https://mistershell.example.com). The worker re-resolves this name on every reconnect, so a health-checked LB/GSLB is the worker’s failover path.
WORKER_TOKENAuthentication token issued when you registered the worker in the UI.

Remote worker — optional

VariableDefaultDescription
LOG_LEVELINFOLogging level.

Guest proxy — required (optional component, Session Proxy add-on)

VariableDescription
MISTERSHELL_URLPublic URL of your core deployment — your load balancer / GSLB hostname. The proxy connects out to it.
PROXY_TOKENAuthentication token issued when you registered the proxy in the UI (starts with prx_).

The proxy is a single web process listening on :8000 (override with PORT); it terminates no TLS of its own — put your own TLS-terminating load balancer in front of it. Both variables are mandatory; the agent exits immediately if either is missing.

IDS sensor — required (optional component, IDS Sensors add-on)

VariableDescription
MISTERSHELL_URLPublic URL of your core deployment — your load balancer / GSLB hostname. The sensor connects out to it.
SENSOR_TOKENAuthentication token issued when you registered the sensor in the UI (starts with sns_).

The sensor captures live traffic, so it needs host networking and packet-capture capabilities (NET_RAW, NET_ADMIN, SYS_NICE). The capture interfaces and home network you set on its registration form are delivered by the core on connect — you do not pass them as environment variables, but the host must actually have those interfaces. Both variables are mandatory; the agent exits immediately if either is missing.

TLS

The core container terminates TLS on :443. On first start it generates a self-signed bootstrap certificate so that the instance is immediately usable; browsers will display the expected untrusted-certificate warning until you replace the certificate.

Three common options to install a real certificate:

Terminate TLS at a cloud load balancer (ALB, NLB, Cloud Load Balancer, Cloudflare, etc.) and forward to the container on :443 — pass-through — or on :80. When terminating at the load balancer, set CORS_ORIGINS only if browsers reach the API from a different origin than the UI.

2. From the web UI

Sign in as an administrator, open Settings → System → Advanced Settings, and set:

  • nginx_tls_certificate_pem — PEM-encoded certificate chain.
  • nginx_tls_private_key_pem — PEM-encoded private key.

Both values are stored encrypted. On a single-core deployment, as soon as both are valid MisterShell writes them to its reverse proxy’s runtime files and reloads the proxy. If either value is invalid, the previous working certificate is kept.

On a multi-core cluster the settings live in the shared database, so you set them once from any core. Restart every core for the new certificate to take effect — one at a time keeps the cluster serving throughout.

3. Mount files at startup

docker run -d -p 443:443 -p 80:80 \
  -e DB_ENCRYPTION_KEY=your-secret-key \
  -v /path/to/server.crt:/etc/ssl/certs/mistershell.crt:ro \
  -v /path/to/server.key:/etc/ssl/private/mistershell.key:ro \
  mistershell/core:latest

A mounted certificate is local to its container and is not shared through the database, so on a multi-core cluster mount the new files on every core and restart each one.

Trusting a private CA

If MisterShell must make outbound TLS connections to services signed by your own certificate authority (LDAP, SMTP, webhooks, AI providers, …), mount the CA certificates at /etc/certs — every image (core, worker, sensor, proxy) picks up *.crt / *.pem files there at startup. See CA Certificates for this and the workspace-wide trust store managed from the UI.

Register a worker

  1. Sign in to the web UI as an administrator.
  2. Open the Fabric page’s Estate tab.
  3. Click Create → Worker, fill in name and location, and save.
  4. Copy the token from the modal — it is only shown once.
  5. On the worker host, pass the token as WORKER_TOKEN to docker run.

Once the worker is running and reaches the core, its status on the Workers tab transitions to online within a few seconds.

Route resources to specific workers

Task routing is location-based: each worker is bound to a location, and a task runs on an online worker in its target location (walking up to a parent location if none is online there). A resource-bound task inherits the location of its resource, so to control which worker handles which resources, assign those resources and the worker to the same location. Set a worker’s location when you register it, or change it later from the Workers tab.

Reverse proxy and X-Forwarded-For

When another reverse proxy or load balancer sits in front of the core, enable client IP forwarding so audit logs and session records capture real client addresses:

  1. Open Settings → System → Advanced Settings.
  2. Find enable_x_forwarded_for_header and toggle it to true.
  3. Save. Changes take effect immediately.

Your proxy must overwrite (not append to) X-Forwarded-For; otherwise end users can choose their own source address.

Scaling

MisterShell collects inventory data (snapshots, health checks) in parallel. Two knobs control throughput:

  • data_collection_concurrency — the maximum number of resources the platform polls in parallel. Tune this from Settings → System → Advanced Settings.
  • DB_POOL_SIZE + DB_MAX_OVERFLOW — how many database connections are available to the core process.

Each parallel dispatch uses one database connection. MisterShell caps parallel dispatches at half of the total pool capacity, reserving the other half for API traffic, task lifecycle, and other services.

Suggested sizing:

Managed resourcesdata_collection_concurrencyDB_POOL_SIZEDB_MAX_OVERFLOWTotal poolEffective parallel
< 1002 (default)20 (default)40 (default)602
100 – 5001020406010
500 – 1 0002030508020
1 000+30406010030

Example for a 1 000-resource deployment:

docker run -d -p 443:443 -p 80:80 \
  -e DB_ENCRYPTION_KEY=your-secret \
  -e DB_POOL_SIZE=40 \
  -e DB_MAX_OVERFLOW=60 \
  mistershell/core:latest

Then raise data_collection_concurrency to 30 in Settings → System → Advanced Settings.

For multi-core topologies, horizontal scaling (more cores) also spreads API and session load — see the HA Cluster page.

Docker Compose examples

Production core

version: '3.8'
services:
  mistershell:
    image: mistershell/core:latest
    ports:
      - "443:443"
      - "80:80"
    environment:
      DB_ENCRYPTION_KEY: ${DB_ENCRYPTION_KEY}
      CORS_ORIGINS: ${CORS_ORIGINS}
    # Required for web-application sessions (headless Chrome sandbox) — see
    # "Web sessions (headless Chrome)" above.
    cap_add:
      - SYS_ADMIN
    security_opt:
      - apparmor:unconfined
    volumes:
      - mistershell_data:/data
    restart: unless-stopped

volumes:
  mistershell_data:

Remote worker

version: '3.8'
services:
  worker:
    image: mistershell/worker:latest
    environment:
      MISTERSHELL_URL: ${MISTERSHELL_URL}
      WORKER_TOKEN: ${WORKER_TOKEN}
    # Required for web-application sessions (headless Chrome sandbox) — see
    # "Web sessions (headless Chrome)" above.
    cap_add:
      - SYS_ADMIN
    security_opt:
      - apparmor:unconfined
    restart: unless-stopped
    # Required when the worker must reach devices on the host network
    network_mode: host

Guest proxy (optional, proxy feature)

Run the proxy on a public-facing host external guests can reach, behind your own TLS-terminating load balancer:

docker run -d -p 8000:8000 \
  -e MISTERSHELL_URL=https://mistershell.example.com \
  -e PROXY_TOKEN=prx_1_your-token \
  --name mistershell-proxy \
  mistershell/proxy:latest
version: '3.8'
services:
  proxy:
    image: mistershell/proxy:latest
    ports:
      - "8000:8000"          # front this with your own TLS-terminating LB
    environment:
      MISTERSHELL_URL: ${MISTERSHELL_URL}
      PROXY_TOKEN: ${PROXY_TOKEN}
    restart: unless-stopped

IDS sensor (optional, ids feature)

The sensor needs host networking and packet-capture capabilities to see traffic:

docker run -d --network host \
  --cap-add NET_RAW --cap-add NET_ADMIN --cap-add SYS_NICE \
  -e MISTERSHELL_URL=https://mistershell.example.com \
  -e SENSOR_TOKEN=sns_1_your-token \
  --name mistershell-sensor \
  mistershell/sensor:latest

Set the capture interfaces and home network on the sensor’s registration form; the sensor pulls them from the core on connect. The host must actually have those interfaces.

Multi-core and per-region Compose / Kubernetes examples live on the HA Cluster and Distributed with HA pages.

Health checks

EndpointDescription
GET https://<core>/health/Core server liveness check.
GET https://<core>/health/readyCore server readiness check — use this as the load-balancer health check in multi-core topologies.
GET https://<core>/health/detailedReadiness plus a best-effort cluster block (node id, leader flag, mesh state, peer count) for cluster diagnostics. Requires authentication (an admin session or API key as a bearer token) — for an unauthenticated view from inside the container, use msh show cluster on the operator console.

On the Workers tab, every registered worker reports its own liveness via heartbeats — online in the table means the worker is reachable, authenticated, and accepting tasks.

Troubleshooting

Collect logs

docker logs mistershell
docker logs mistershell-worker

Common issues

SymptomLikely cause and remedy
Worker stays in offline stateVerify that WORKER_TOKEN matches the one shown at creation time and that the worker host can reach wss://<core>/nats through your proxy.
Wrong client IPs in session and audit logsTurn on enable_x_forwarded_for_header under Settings → System → Advanced Settings only when another reverse proxy in front of the core overwrites (not appends to) X-Forwarded-For.
Web-application sessions fail to start (SSH/RDP/VNC still work)The headless-Chrome sandbox is blocked by the container’s default confinement. Run the core/worker container with --cap-add SYS_ADMIN --security-opt apparmor=unconfined. Logs show Failed to move to new namespace: Operation not permitted. See Web sessions (headless Chrome).
Core fails to start with “Cannot connect to database”Check DATABASE_URL, reachability of the database host, and credentials. For the embedded variant, confirm the data volume is writable.
Multi-core: a core refuses to start, complaining that DATABASE_URL or APP_REDIS_URL is localA core with a non-localhost THIS_CORE_URL requires external DATABASE_URL and APP_REDIS_URL (shared coherence-sensitive state). Point both at your managed services.
Multi-core: cores don’t form a mesh / no leader is electedConfirm every core shares the same THIS_CORE_REGION, that the cluster route port 6222 (and gateway port 7222 across regions) is open between cores, and that NODE_ID is unique and stable per core. Inspect msh show cluster inside a core container, or GET /health/detailed (authenticated).
WebSocket disconnects shortly after connectingYour reverse proxy or load balancer must allow WebSocket upgrades and set a generous idle-timeout (ideally 5+ minutes).
Browser warns that the certificate is not trustedThe default bootstrap certificate is self-signed. Replace it with a trusted certificate using any of the TLS methods above.

What’s next