Monitor → Metrics (Prometheus)
Every Core serves platform self-metrics at https://<core-address>/metrics in the standard Prometheus text format, ready to scrape into your existing Prometheus, Grafana Agent, or any OpenMetrics-compatible collector. The endpoint reports both this Core’s node-local activity (message-bus traffic, stream depths) and fleet-wide state (components, tasks, sessions) — so scraping a single Core behind a load balancer gives you the full picture.
Set up a scrape credential
The endpoint requires a bearer token tied to a user with the Fabric read permission.
- Open Settings → Users & Roles, and on the Users tab click Create User. Give the service account a recognizable name, for example
svc-prometheus. - On the Roles tab, create a role (for example
metrics-reader) and grant it onlyapp.fabric.read. Assign the role to the service user. - Sign in as the service user, open My Account → Account Security → My API key, and generate the account’s API key. Copy the key — it is shown only once.
Configure Prometheus
scrape_configs:
- job_name: mistershell
scheme: https
metrics_path: /metrics
bearer_token: "<the service user's API key>"
static_configs:
- targets: ["core-1.example.com:443", "core-2.example.com:443"]
# If your Cores use a private CA or self-signed certificate:
tls_config:
insecure_skip_verify: true
A 15–60 second scrape interval is appropriate. Scrapes are lightweight, are never rate-limited, and are not recorded in the audit log.
Scraping more than one Core
Fleet-wide metrics (components, tasks, sessions, worker slots) are read from the shared database and are identical from every Core — Prometheus’s instance label tells the sources apart. When you scrape several Cores, deduplicate fleet totals in your queries:
max without(instance) (mistershell_tasks_stuck)
Node-local metrics (mistershell_bus_*, mistershell_stream_*, the process_* families) genuinely differ per Core — aggregate or graph them per instance.
Metric reference
All platform metrics are prefixed mistershell_. Standard Python process metrics (process_*, python_*) are included as well.
| Metric | Type | What it tells you |
|---|---|---|
mistershell_info{version,node_id,region} | gauge (1) | The Core’s version, node id, and region. |
mistershell_cluster_is_leader | gauge (0/1) | Whether this Core currently coordinates the cluster. |
mistershell_cluster_leadership_epoch | gauge | Leadership generation counter (present on the leader only). |
mistershell_cluster_peers_total | gauge | Registered Cores, including this one. |
mistershell_cluster_peer_info{node_id,region,version} | gauge (1) | One series per registered Core. |
mistershell_bus_messages_in_total / ..._out_total | counter | Messages received/sent by this Core’s message bus. Use rate(). |
mistershell_bus_bytes_in_total / ..._out_total | counter | Bytes received/sent by this Core’s message bus. Use rate(). |
mistershell_bus_connections{class} | gauge | Bus connections on this Core by class: client, worker, sensor, proxy, unclassified. |
mistershell_bus_routes_up | gauge | Established links from this Core to its same-region peers. |
mistershell_stream_messages{stream} | gauge | Messages currently retained per delivery stream on this Core. |
mistershell_stream_bytes{stream} | gauge | Bytes currently retained per delivery stream. |
mistershell_stream_consumers{stream} | gauge | Consumers attached per delivery stream. |
mistershell_components{role,status} | gauge | Fleet component counts by role (worker/sensor/proxy) and status. |
mistershell_component_up{role,name} | gauge (0/1) | Per-component online flag — alert when a specific worker goes down. |
mistershell_component_info{role,name,version} | gauge (1) | One series per fleet component with its reported version. |
mistershell_worker_slots_used / mistershell_worker_slots_total | gauge | Task slots in use vs. capacity across online workers. |
mistershell_tasks{status} | gauge | Task counts by status (pending, queued, running, success, failed, …). |
mistershell_tasks_stuck | gauge | Tasks that exceeded their execution or recovery window. Alert on > 0. |
mistershell_sessions_active{type} | gauge | Active interactive sessions by type (ssh, rdp, vnc, web, …). |
mistershell_dependency_up{dependency} | gauge (0/1) | Backend dependency checks, one series per check. Label values: database (the platform database), redis (the shared cache), nats (the platform message bus), graphical_recording_stream (the capture stream for graphical session recording), critical_consumers (the platform’s critical event consumers), licensing_reconciliation (license entitlement reconciliation), licensing_ha (flags a multi-Core deployment running without the multi-Core entitlement), disk_space (local disk capacity). |
mistershell_dependency_latency_seconds{dependency} | gauge | Latency of each dependency check. |
Useful alerts
# A specific worker went offline
mistershell_component_up{role="worker"} == 0
# Stuck tasks anywhere in the fleet
max without(instance) (mistershell_tasks_stuck) > 0
# A backend dependency is failing on some Core
mistershell_dependency_up == 0
# Worker capacity saturated
max without(instance) (mistershell_worker_slots_used)
>= max without(instance) (mistershell_worker_slots_total)
Behaviour under partial failure
A scrape always answers with whatever can be gathered — it never fails outright because one subsystem is down:
- If a backend dependency is unreachable, its
mistershell_dependency_upreports0and the metric families that depend on it are absent from that scrape (treat absence as “unknown”, not zero). - The message-bus and stream families come from a periodic on-node sample. Right after a Core starts (or if sampling is interrupted) they are omitted until the next sample lands — usually within seconds.
Permissions
- Scrape
/metrics:app.fabric.read(via the bearer token’s user).