Deployment
MisterShell is distributed as four published container images. You pull them directly from the registry.
| Image | Purpose |
|---|---|
mistershell/core:latest | The control plane — web UI, API, event bus, session gateway, and an embedded worker. This is what your users connect to. |
mistershell/worker:latest | A remote worker for distributed deployments. Install one per site, region, or network segment. |
mistershell/proxy:latest | A public-facing session proxy that lets external guests reach a shared session without exposing your core. Optional; requires the Session Proxy add-on — see Proxies. |
mistershell/sensor:latest | A remote intrusion-detection probe that passively watches network traffic and raises alerts. Optional; requires the IDS Sensors add-on — see Sensors. |
Every deployment uses core (and, for distributed topologies, worker). The proxy and sensor are optional add-ons you install only if you use external session sharing or intrusion detection, and each needs its own license add-on.
Pin every image to the chosen release tag (mistershell/core:<version>, for example) in production so that upgrades are deliberate.
This page is the shared reference for every deployment: prerequisites, environment variables, TLS, web sessions, worker registration, scaling, and troubleshooting. Pick your topology below — each topology page has the components, boundaries, diagram, benefits, and step-by-step guidance specific to it, and links back here for the shared bits.
Choose a deployment topology
MisterShell supports four topologies. They differ along two axes: how many cores per region (redundancy) and how many regions (geographic reach).
| Topology | Cores | Regions | In-region HA | Geographic | External services you run |
|---|---|---|---|---|---|
| All in One | 1 | 1 | — | — | None (all embedded) |
| HA Cluster | ≥3 | 1 | ✅ | — | Database, App Redis, LB, object store |
| Distributed | 1 / region | ≥2 | — | ✅ | Central database, App Redis, GSLB, object store |
| Distributed with HA | ≥3 / region | ≥2 | ✅ | ✅ | Central database, App Redis, GSLB + per-region LBs, object store |
- All in One — one container, everything embedded. Evaluation, proofs-of-concept, small single-zone estates.
- HA Cluster — multiple cores in one region, no single point of failure in the application tier.
- Distributed — one core per region for geographic locality, without per-region redundancy.
- Distributed with HA — redundant cores in each of several regions; the full topology.
A common boundary applies to every multi-core topology: MisterShell makes its own application tier highly available (leader election, mesh, replicated recordings), but it depends on you to run the load balancer / GSLB, the database, the app-state store (App Redis), and the object store — and to make those highly available. Note that cluster leadership itself is coordinated through the database (a lease that a surviving core takes over automatically), so the database’s availability is what keeps failover working. Each topology page spells out exactly what falls on which side of that line.
Upgrades and rollback
Back up the MisterShell database and verify that the backup is readable before
every upgrade. Also preserve DB_ENCRYPTION_KEY and any externally managed
recording store according to your retention policy. A successful application
startup or schema migration is not a substitute for a restorable backup.
See Database Backup and Restore for the scheduled
task, MSH commands, and the different single-Core and multi-Core procedures.
For a single-Core deployment, back up the data, replace the container with the new image, and start it. Pending database migrations and registry synchronization run automatically at startup.
Cold and rolling upgrade considerations apply only to multi-Core deployments. Within one major version, eligible clusters support rolling upgrades unless the release notes require a coordinated cold upgrade. For a cold upgrade, stop new sessions, let active sessions finish, stop every old Core and Worker, take the verified backup, and then start only the new version. Multi-Core upgrades across major versions require this coordinated procedure; old and new major versions must not run concurrently.
Rollback after a major-version upgrade means restoring the pre-upgrade database backup and the matching previous-version containers together. Do not use schema downgrades as a data-recovery procedure. Confirm database, recording-store, and encryption-key recovery in a non-production environment before relying on the procedure.
Prerequisites
- A Linux host with a container runtime (Docker, Podman, or equivalent) per core and per remote worker.
- DNS / TLS planning: the core server answers on
:443(HTTPS + WebSocket upgrades) and:80. By default:80serves the full application over plain HTTP (for use behind your own TLS-terminating reverse proxy); setENABLE_HTTP_REDIRECT=trueto make:80serve only health checks and HTTP → HTTPS redirects. Do not expose:80to untrusted networks unless the redirect is enabled. - If you enable the syslog collector on the core’s embedded worker, also map the collector listener port (
514/udpand514/tcpby default). - Native SSH access (
ssh/sftp/scpagainst sessions and the file catalog) listens on thesshgw_portsetting (default2222); map it 1:1 at the host, in Compose, and on the Kubernetes load balancer if you want operators to use it. The gateway is started and supervised by the Core itself — no separate container or process to run. - For multi-site deployments, outbound HTTPS/WebSocket connectivity from each worker host back to the core’s public URL.
- For any multi-core topology, an externally managed relational database and app-state store (App Redis) reachable from every core, plus a load balancer / GSLB in front of the cores. See the topology pages for the exact requirements.
Web sessions (headless Chrome)
Web-application sessions open the target site in a headless Chrome running next to the worker. Chrome is already bundled in both the core and worker images (including the embedded worker inside the core image), so there is nothing to install. Chrome keeps its renderer sandbox on — the isolation boundary between the visited page and the worker — and that sandbox needs to create namespaces, which the container blocks by default.
Requirement — run the core container (and any remote worker that hosts web sessions) with these two settings:
--cap-add SYS_ADMIN --security-opt apparmor=unconfined
--cap-add SYS_ADMINlets the container create the namespaces Chrome’s sandbox needs, while keeping the default seccomp filter on.--security-opt apparmor=unconfinedclears Ubuntu 22.04/24.04’s AppArmor user-namespace restriction; it’s a harmless no-op on hosts without AppArmor, so the same two settings work everywhere.
Without them, web sessions fail to start and the logs show Failed to move to new namespace: Operation not permitted. Other session types (SSH, RDP, VNC, cloud,
DB) are unaffected. Don’t disable the sandbox with Chrome --no-sandbox — that
removes the renderer isolation and is a security regression.
Kubernetes — per core/worker pod securityContext:
securityContext:
capabilities:
add: ["SYS_ADMIN"]
# AppArmor: set the container annotation to unconfined, e.g.
# container.apparmor.security.beta.kubernetes.io/<container>: unconfined
Portainer — deploy the core/worker as a Stack using the Compose YAML in
Docker Compose examples; the cap_add / security_opt
keys apply as-is. Edit and redeploy the stack to recreate the container with
the options (changing them on a running container has no effect). On a Portainer
Swarm environment security_opt is ignored, so rely on cap_add: SYS_ADMIN.
Notes:
- A simpler but broader alternative is
--security-opt seccomp=unconfined --security-opt apparmor=unconfined; it works too, but turns off the whole seccomp filter — prefer theSYS_ADMINform, which keeps it on. - Locked-down platforms (Kubernetes restricted/baseline PodSecurity, OpenShift, Docker Swarm, rootless/gVisor) may forbid these options; there, web sessions can’t run and every other feature still works.
Environment variables
Core container — required
| Variable | Description |
|---|---|
DB_ENCRYPTION_KEY | Secret used to encrypt sensitive values stored in the database. Treat this as a credential and back it up. In a multi-core deployment, every core uses the same value. |
Core container — optional
| Variable | Default | Description |
|---|---|---|
DATABASE_URL | embedded database | Connection string for the relational database. A localhost / 127.0.0.1 host activates the embedded database; an external host routes to your own service. Mandatory (external) in any multi-core topology. |
DB_POOL_SIZE | 20 | Persistent database connections. Scale alongside data_collection_concurrency — see Scaling. |
DB_MAX_OVERFLOW | 40 | Burst database connections under load. Total pool capacity = DB_POOL_SIZE + DB_MAX_OVERFLOW. |
CACHE_REDIS_URL | embedded cache | Connection string for the high-performance in-memory hot cache. Same localhost rule as DATABASE_URL. May stay per-core even in a cluster. |
APP_REDIS_URL | embedded app store | Connection string for the app-state store (live sessions, locks, cluster coordination state). Same localhost rule as DATABASE_URL. Mandatory (external, shared) in any multi-core topology, and the external instance must use maxmemory-policy noeviction. |
CORS_ORIGINS | (empty) | Comma-separated list of browser origins allowed for cross-origin API requests. Set this only when the UI and API are served from different origins. |
LOG_LEVEL | INFO | One of DEBUG, INFO, WARNING, ERROR, CRITICAL. |
TLS_CERT_CN | localhost | Common name used for the self-signed bootstrap certificate generated on first start. |
ENABLE_HTTP_REDIRECT | false | When true, :80 serves only health checks and redirects everything else to HTTPS. When false (the default), :80 serves the full application over plain HTTP — intended for running behind your own TLS-terminating reverse proxy. |
EMBEDDED_WORKER | true (single-node) | Start the embedded worker inside the core container. In multi-node mode it defaults off and cannot be turned on — setting it to true there is a startup error; run remote workers instead. |
Core container — clustering (multi-core topologies only)
These select and shape a multi-core deployment. On a single-core (All in One) deployment, leave them at their defaults.
| Variable | Default | Description |
|---|---|---|
THIS_CORE_URL | http://localhost:8000 | This core’s inter-core address — how peer cores reach it on the cluster route port (:6222) within a region and the gateway port (:7222) across regions. A non-localhost value is what switches the core into multi-node mode. This is not the address workers use — that is MISTERSHELL_URL, which points at your load balancer. |
THIS_CORE_REGION | default | Region name; also the cluster name. Same on every core in a region. Must match ^[a-z0-9-]+$ (≤32 chars). |
NODE_ID | derived from THIS_CORE_URL | Stable, unique lease-holder identity (used for leader election and the cluster mesh). If unset, derived from the THIS_CORE_URL host and port (e.g. http://core-0.internal:8000 → core-0.internal_8000, unsafe characters mapped to _); set explicitly on ordered platforms (e.g. from a StatefulSet pod name). Must match ^[A-Za-z0-9._-]+$ (1–64 chars). |
NATS_STREAM_REPLICAS | region size | Replica count for the region’s event/recording streams (1–9). Defaults to the number of cores in the region, so you normally leave it unset. |
Remote worker — required
| Variable | Description |
|---|---|
MISTERSHELL_URL | Public URL of your core deployment — your load balancer / GSLB hostname (for example https://mistershell.example.com). The worker re-resolves this name on every reconnect, so a health-checked LB/GSLB is the worker’s failover path. |
WORKER_TOKEN | Authentication token issued when you registered the worker in the UI. |
Remote worker — optional
| Variable | Default | Description |
|---|---|---|
LOG_LEVEL | INFO | Logging level. |
Guest proxy — required (optional component, Session Proxy add-on)
| Variable | Description |
|---|---|
MISTERSHELL_URL | Public URL of your core deployment — your load balancer / GSLB hostname. The proxy connects out to it. |
PROXY_TOKEN | Authentication token issued when you registered the proxy in the UI (starts with prx_). |
The proxy is a single web process listening on :8000 (override with PORT); it terminates no TLS of its own — put your own TLS-terminating load balancer in front of it. Both variables are mandatory; the agent exits immediately if either is missing.
IDS sensor — required (optional component, IDS Sensors add-on)
| Variable | Description |
|---|---|
MISTERSHELL_URL | Public URL of your core deployment — your load balancer / GSLB hostname. The sensor connects out to it. |
SENSOR_TOKEN | Authentication token issued when you registered the sensor in the UI (starts with sns_). |
The sensor captures live traffic, so it needs host networking and packet-capture capabilities (NET_RAW, NET_ADMIN, SYS_NICE). The capture interfaces and home network you set on its registration form are delivered by the core on connect — you do not pass them as environment variables, but the host must actually have those interfaces. Both variables are mandatory; the agent exits immediately if either is missing.
TLS
The core container terminates TLS on :443. On first start it generates a self-signed bootstrap certificate so that the instance is immediately usable; browsers will display the expected untrusted-certificate warning until you replace the certificate.
Three common options to install a real certificate:
1. Behind a load balancer (recommended)
Terminate TLS at a cloud load balancer (ALB, NLB, Cloud Load Balancer, Cloudflare, etc.) and forward to the container on :443 — pass-through — or on :80. When terminating at the load balancer, set CORS_ORIGINS only if browsers reach the API from a different origin than the UI.
2. From the web UI
Sign in as an administrator, open Settings → System → Advanced Settings, and set:
nginx_tls_certificate_pem— PEM-encoded certificate chain.nginx_tls_private_key_pem— PEM-encoded private key.
Both values are stored encrypted. On a single-core deployment, as soon as both are valid MisterShell writes them to its reverse proxy’s runtime files and reloads the proxy. If either value is invalid, the previous working certificate is kept.
On a multi-core cluster the settings live in the shared database, so you set them once from any core. Restart every core for the new certificate to take effect — one at a time keeps the cluster serving throughout.
3. Mount files at startup
docker run -d -p 443:443 -p 80:80 \
-e DB_ENCRYPTION_KEY=your-secret-key \
-v /path/to/server.crt:/etc/ssl/certs/mistershell.crt:ro \
-v /path/to/server.key:/etc/ssl/private/mistershell.key:ro \
mistershell/core:latest
A mounted certificate is local to its container and is not shared through the database, so on a multi-core cluster mount the new files on every core and restart each one.
Trusting a private CA
If MisterShell must make outbound TLS connections to services signed by your own certificate authority (LDAP, SMTP, webhooks, AI providers, …), mount the CA certificates at /etc/certs — every image (core, worker, sensor, proxy) picks up *.crt / *.pem files there at startup. See CA Certificates for this and the workspace-wide trust store managed from the UI.
Register a worker
- Sign in to the web UI as an administrator.
- Open the Fabric page’s Estate tab.
- Click Create → Worker, fill in name and location, and save.
- Copy the token from the modal — it is only shown once.
- On the worker host, pass the token as
WORKER_TOKENtodocker run.
Once the worker is running and reaches the core, its status on the Workers tab transitions to online within a few seconds.
Route resources to specific workers
Task routing is location-based: each worker is bound to a location, and a task runs on an online worker in its target location (walking up to a parent location if none is online there). A resource-bound task inherits the location of its resource, so to control which worker handles which resources, assign those resources and the worker to the same location. Set a worker’s location when you register it, or change it later from the Workers tab.
Reverse proxy and X-Forwarded-For
When another reverse proxy or load balancer sits in front of the core, enable client IP forwarding so audit logs and session records capture real client addresses:
- Open Settings → System → Advanced Settings.
- Find
enable_x_forwarded_for_headerand toggle it totrue. - Save. Changes take effect immediately.
Your proxy must overwrite (not append to) X-Forwarded-For; otherwise end users can choose their own source address.
Scaling
MisterShell collects inventory data (snapshots, health checks) in parallel. Two knobs control throughput:
data_collection_concurrency— the maximum number of resources the platform polls in parallel. Tune this from Settings → System → Advanced Settings.DB_POOL_SIZE+DB_MAX_OVERFLOW— how many database connections are available to the core process.
Each parallel dispatch uses one database connection. MisterShell caps parallel dispatches at half of the total pool capacity, reserving the other half for API traffic, task lifecycle, and other services.
Suggested sizing:
| Managed resources | data_collection_concurrency | DB_POOL_SIZE | DB_MAX_OVERFLOW | Total pool | Effective parallel |
|---|---|---|---|---|---|
| < 100 | 2 (default) | 20 (default) | 40 (default) | 60 | 2 |
| 100 – 500 | 10 | 20 | 40 | 60 | 10 |
| 500 – 1 000 | 20 | 30 | 50 | 80 | 20 |
| 1 000+ | 30 | 40 | 60 | 100 | 30 |
Example for a 1 000-resource deployment:
docker run -d -p 443:443 -p 80:80 \
-e DB_ENCRYPTION_KEY=your-secret \
-e DB_POOL_SIZE=40 \
-e DB_MAX_OVERFLOW=60 \
mistershell/core:latest
Then raise data_collection_concurrency to 30 in Settings → System → Advanced Settings.
For multi-core topologies, horizontal scaling (more cores) also spreads API and session load — see the HA Cluster page.
Docker Compose examples
Production core
version: '3.8'
services:
mistershell:
image: mistershell/core:latest
ports:
- "443:443"
- "80:80"
environment:
DB_ENCRYPTION_KEY: ${DB_ENCRYPTION_KEY}
CORS_ORIGINS: ${CORS_ORIGINS}
# Required for web-application sessions (headless Chrome sandbox) — see
# "Web sessions (headless Chrome)" above.
cap_add:
- SYS_ADMIN
security_opt:
- apparmor:unconfined
volumes:
- mistershell_data:/data
restart: unless-stopped
volumes:
mistershell_data:
Remote worker
version: '3.8'
services:
worker:
image: mistershell/worker:latest
environment:
MISTERSHELL_URL: ${MISTERSHELL_URL}
WORKER_TOKEN: ${WORKER_TOKEN}
# Required for web-application sessions (headless Chrome sandbox) — see
# "Web sessions (headless Chrome)" above.
cap_add:
- SYS_ADMIN
security_opt:
- apparmor:unconfined
restart: unless-stopped
# Required when the worker must reach devices on the host network
network_mode: host
Guest proxy (optional, proxy feature)
Run the proxy on a public-facing host external guests can reach, behind your own TLS-terminating load balancer:
docker run -d -p 8000:8000 \
-e MISTERSHELL_URL=https://mistershell.example.com \
-e PROXY_TOKEN=prx_1_your-token \
--name mistershell-proxy \
mistershell/proxy:latest
version: '3.8'
services:
proxy:
image: mistershell/proxy:latest
ports:
- "8000:8000" # front this with your own TLS-terminating LB
environment:
MISTERSHELL_URL: ${MISTERSHELL_URL}
PROXY_TOKEN: ${PROXY_TOKEN}
restart: unless-stopped
IDS sensor (optional, ids feature)
The sensor needs host networking and packet-capture capabilities to see traffic:
docker run -d --network host \
--cap-add NET_RAW --cap-add NET_ADMIN --cap-add SYS_NICE \
-e MISTERSHELL_URL=https://mistershell.example.com \
-e SENSOR_TOKEN=sns_1_your-token \
--name mistershell-sensor \
mistershell/sensor:latest
Set the capture interfaces and home network on the sensor’s registration form; the sensor pulls them from the core on connect. The host must actually have those interfaces.
Multi-core and per-region Compose / Kubernetes examples live on the HA Cluster and Distributed with HA pages.
Health checks
| Endpoint | Description |
|---|---|
GET https://<core>/health/ | Core server liveness check. |
GET https://<core>/health/ready | Core server readiness check — use this as the load-balancer health check in multi-core topologies. |
GET https://<core>/health/detailed | Readiness plus a best-effort cluster block (node id, leader flag, mesh state, peer count) for cluster diagnostics. Requires authentication (an admin session or API key as a bearer token) — for an unauthenticated view from inside the container, use msh show cluster on the operator console. |
On the Workers tab, every registered worker reports its own liveness via heartbeats — online in the table means the worker is reachable, authenticated, and accepting tasks.
Troubleshooting
Collect logs
docker logs mistershell
docker logs mistershell-worker
Common issues
| Symptom | Likely cause and remedy |
|---|---|
| Worker stays in offline state | Verify that WORKER_TOKEN matches the one shown at creation time and that the worker host can reach wss://<core>/nats through your proxy. |
| Wrong client IPs in session and audit logs | Turn on enable_x_forwarded_for_header under Settings → System → Advanced Settings only when another reverse proxy in front of the core overwrites (not appends to) X-Forwarded-For. |
| Web-application sessions fail to start (SSH/RDP/VNC still work) | The headless-Chrome sandbox is blocked by the container’s default confinement. Run the core/worker container with --cap-add SYS_ADMIN --security-opt apparmor=unconfined. Logs show Failed to move to new namespace: Operation not permitted. See Web sessions (headless Chrome). |
| Core fails to start with “Cannot connect to database” | Check DATABASE_URL, reachability of the database host, and credentials. For the embedded variant, confirm the data volume is writable. |
Multi-core: a core refuses to start, complaining that DATABASE_URL or APP_REDIS_URL is local | A core with a non-localhost THIS_CORE_URL requires external DATABASE_URL and APP_REDIS_URL (shared coherence-sensitive state). Point both at your managed services. |
| Multi-core: cores don’t form a mesh / no leader is elected | Confirm every core shares the same THIS_CORE_REGION, that the cluster route port 6222 (and gateway port 7222 across regions) is open between cores, and that NODE_ID is unique and stable per core. Inspect msh show cluster inside a core container, or GET /health/detailed (authenticated). |
| WebSocket disconnects shortly after connecting | Your reverse proxy or load balancer must allow WebSocket upgrades and set a generous idle-timeout (ideally 5+ minutes). |
| Browser warns that the certificate is not trusted | The default bootstrap certificate is self-signed. Replace it with a trusted certificate using any of the TLS methods above. |
What’s next
- Pick and stand up your deployment topology.
- Sign in for the first time and walk through Settings → System → Config to tune email, backups, TLS, data retention, and performance.
- Register and install any remote workers your environment requires.
- Install a license from Settings → System → Licensing to go beyond the Free edition’s 25 resources and unlock licensed features.
- Onboard your first resource from Manage → Browse.