Skip to content
User Guide

Operator Console (msh)

Every MisterShell container — Core, Worker, Sensor, and Proxy — ships an operator console for diagnostics, recovery, and support bundles. It works even when the web interface or API is unavailable, because on a Core it talks to the platform’s data stores directly from inside the container.

Open a shell in any MisterShell container and type:

msh

You get an interactive prompt in the style of a network-device CLI:

  • ? at any cursor position lists the valid next words with one-line help, without submitting the line.
  • Tab completes the current word.
  • Abbreviations work everywhere a prefix is unique: sh clu runs show cluster, sh tech runs show tech-support.
  • exit or quit leaves the console. Command history is kept under /data and survives restarts.

What each role can do

The console adapts to the container it runs in. Every role offers the host-level and support commands — show version, show system, show log, show tech-support, write tech-support, and bash. The Core adds the platform-wide inspection and recovery commands below (cluster, streams, fleet, tasks, sessions, settings, licensing, the clear … recovery verbs, reset admin-user, license install/remove, sync registries, and migrate …, and database backup/restore), because only the Core is connected to the platform’s data stores. On a Worker, Sensor, or Proxy the console is a local diagnostic tool; on a Core it is also a recovery tool.

One-shot mode

For scripts and runbooks, pass the command as arguments and the console runs it and exits:

msh show health
msh -y clear tasks stuck

The exit code is 0 on success and 1 on any error. Commands that change state require confirmation: interactively they ask ([confirm] — press Enter to proceed); in one-shot mode pass -y (or --yes) before the command, or set the environment variable MSH_AUTO_APPROVE=true.

Host commands (every role)

CommandShows
show versionVersion, role, and uptime of this container; on a Core also its node id, region, and advertised URL
show systemHost summary: load, CPU, memory, swap, disk for / and /data, process count
show system cpuLoad, per-core usage, and the top processes by CPU
show system memoryMemory and swap totals plus the top processes by memory
show system diskUsage per mounted filesystem
show system networkPer-interface received/sent bytes and error counts
show log [last <N>] [-f]Tail the container’s log file (default last 100 lines); -f (or follow) streams new lines until you press Ctrl-C. The log file lives under /data, so on a Worker, Sensor, or Proxy started without a /data volume there is nothing to tail
bashDrop to an interactive shell and return to msh on exit

Core inspection commands

These are available on a Core.

CommandShows
show healthDatabase / cache / message-delivery / disk checks with latency
show clusterCurrent leader, leadership epoch, this node’s role
show cluster peersEvery core node: id, region, version, address, age
show streamsInternal delivery streams: messages, size, consumers, rates
show stream <name>One stream in detail, including consumers and backlog
show workers [detail]Worker status, version, heartbeat age, slots used/max
show sensorsSensor fleet status
show proxiesProxy fleet status, including active sessions
show tasks [stuck|queued|running]Task summary by status; filters list rows
show task <id>One task: its chain, status, worker, and timings
show sessionsActive sessions by type, with per-session rows
show settings [<key>]Current vs default values (secrets always masked)
show licenseInstalled licenses and capacity vs current usage
show backupsFull database archives on the configured SFTP target

Long tables are capped and end with (N more — use filters).

Recovery commands (Core)

Every recovery command shows its scope, asks for confirmation, and writes an audit-log entry (actor console) recording the exact command line and the affected items.

CommandRecovers
clear task <id>Force a stuck task into a terminal state and free its worker slot
clear tasks stuckSweep: fail everything stuck-detection finds
clear sessions orphanedDelete leftover session state whose session no longer exists (live sessions are never touched)
clear lock scheduler <name|all>Release a wedged background-job lock
clear presence <worker|sensor|proxy> <id>Drop a stale presence/statistics entry
clear cache settingsForce settings to be re-read from the database
clear cache permissionsForce user permissions to be re-evaluated

Reset the local admin

reset admin-user (Core) recreates the built-in local administrator when you are locked out — for example after the last admin password is lost. It restores the fixed admin identity, generates a fresh password and prints it once, clears any account lockout, and records the action in the audit log. Copy the printed password immediately and change it after you sign in.

Maintenance commands (Core)

Like the recovery verbs, these ask for confirmation and are recorded in the audit log.

CommandDoes
license installInstall a license from the console when the UI is unreachable. The license is read from stdin only (msh license install < license.jws), previewed, and confirmed before it is stored
license remove <license-uuid>Remove a stored license by id, with the same preview + confirm
sync registriesRe-seed the platform registries from the running code (settings, scheduled tasks, metrics, snaps, facts, …) on a live cluster — idempotent, existing values are preserved. Run after a rolling upgrade to synchronize the platform catalog without a cold restart
migrate rdk-database-typesCold-upgrade preflight: discover and save the native type of MySQL/MariaDB resources whose type is unresolved. The update is atomic: if any resource cannot be resolved, no rows are changed
backup databaseCreate, validate, and upload a full PostgreSQL custom archive using the configured SFTP settings
restore database file <path>Validate and restore a local archive; single-Core deployments only
restore database sftp <filename>Download to private staging, then use the same guarded restore pipeline; single-Core deployments only

Database restore replaces the configured embedded or external PostgreSQL database and stops the Core after success. Multi-Core restore must be orchestrated outside MisterShell after every Core is stopped. See Database Backup and Restore before using either restore command.

Tech-support report

One report, two verbs, available on every role:

  • show tech-support prints the complete diagnostic report to the console — every host command’s output plus, on a Core, the inspection commands, a redacted environment listing, disk usage of /data, rendered service configuration, and a recent log tail when a log file exists in the container. On a Worker, Sensor, or Proxy the report contains the sections that role can answer (version, system, and log).
  • write tech-support saves the same report as /data/tech-support-<node-id>-<timestamp>.tar.gz and prints the path — attach this file to a support case.

Secrets (credential values, keys, tokens, passwords) are always masked in both forms.

Permissions

  • None inside the platform’s permission model — access to a container’s console is the trust boundary. On a Core, recovery and reset actions are recorded in the audit log.