Skip to content
User Guide

Monitor → Metrics (Prometheus)

Every Core serves platform self-metrics at https://<core-address>/metrics in the standard Prometheus text format, ready to scrape into your existing Prometheus, Grafana Agent, or any OpenMetrics-compatible collector. The endpoint reports both this Core’s node-local activity (message-bus traffic, stream depths) and fleet-wide state (components, tasks, sessions) — so scraping a single Core behind a load balancer gives you the full picture.

Set up a scrape credential

The endpoint requires a bearer token tied to a user with the Fabric read permission.

  1. Open Settings → Users & Roles, and on the Users tab click Create User. Give the service account a recognizable name, for example svc-prometheus.
  2. On the Roles tab, create a role (for example metrics-reader) and grant it only app.fabric.read. Assign the role to the service user.
  3. Sign in as the service user, open My Account → Account Security → My API key, and generate the account’s API key. Copy the key — it is shown only once.

Configure Prometheus

scrape_configs:
  - job_name: mistershell
    scheme: https
    metrics_path: /metrics
    bearer_token: "<the service user's API key>"
    static_configs:
      - targets: ["core-1.example.com:443", "core-2.example.com:443"]
    # If your Cores use a private CA or self-signed certificate:
    tls_config:
      insecure_skip_verify: true

A 15–60 second scrape interval is appropriate. Scrapes are lightweight, are never rate-limited, and are not recorded in the audit log.

Scraping more than one Core

Fleet-wide metrics (components, tasks, sessions, worker slots) are read from the shared database and are identical from every Core — Prometheus’s instance label tells the sources apart. When you scrape several Cores, deduplicate fleet totals in your queries:

max without(instance) (mistershell_tasks_stuck)

Node-local metrics (mistershell_bus_*, mistershell_stream_*, the process_* families) genuinely differ per Core — aggregate or graph them per instance.

Metric reference

All platform metrics are prefixed mistershell_. Standard Python process metrics (process_*, python_*) are included as well.

MetricTypeWhat it tells you
mistershell_info{version,node_id,region}gauge (1)The Core’s version, node id, and region.
mistershell_cluster_is_leadergauge (0/1)Whether this Core currently coordinates the cluster.
mistershell_cluster_leadership_epochgaugeLeadership generation counter (present on the leader only).
mistershell_cluster_peers_totalgaugeRegistered Cores, including this one.
mistershell_cluster_peer_info{node_id,region,version}gauge (1)One series per registered Core.
mistershell_bus_messages_in_total / ..._out_totalcounterMessages received/sent by this Core’s message bus. Use rate().
mistershell_bus_bytes_in_total / ..._out_totalcounterBytes received/sent by this Core’s message bus. Use rate().
mistershell_bus_connections{class}gaugeBus connections on this Core by class: client, worker, sensor, proxy, unclassified.
mistershell_bus_routes_upgaugeEstablished links from this Core to its same-region peers.
mistershell_stream_messages{stream}gaugeMessages currently retained per delivery stream on this Core.
mistershell_stream_bytes{stream}gaugeBytes currently retained per delivery stream.
mistershell_stream_consumers{stream}gaugeConsumers attached per delivery stream.
mistershell_components{role,status}gaugeFleet component counts by role (worker/sensor/proxy) and status.
mistershell_component_up{role,name}gauge (0/1)Per-component online flag — alert when a specific worker goes down.
mistershell_component_info{role,name,version}gauge (1)One series per fleet component with its reported version.
mistershell_worker_slots_used / mistershell_worker_slots_totalgaugeTask slots in use vs. capacity across online workers.
mistershell_tasks{status}gaugeTask counts by status (pending, queued, running, success, failed, …).
mistershell_tasks_stuckgaugeTasks that exceeded their execution or recovery window. Alert on > 0.
mistershell_sessions_active{type}gaugeActive interactive sessions by type (ssh, rdp, vnc, web, …).
mistershell_dependency_up{dependency}gauge (0/1)Backend dependency checks, one series per check. Label values: database (the platform database), redis (the shared cache), nats (the platform message bus), graphical_recording_stream (the capture stream for graphical session recording), critical_consumers (the platform’s critical event consumers), licensing_reconciliation (license entitlement reconciliation), licensing_ha (flags a multi-Core deployment running without the multi-Core entitlement), disk_space (local disk capacity).
mistershell_dependency_latency_seconds{dependency}gaugeLatency of each dependency check.

Useful alerts

# A specific worker went offline
mistershell_component_up{role="worker"} == 0

# Stuck tasks anywhere in the fleet
max without(instance) (mistershell_tasks_stuck) > 0

# A backend dependency is failing on some Core
mistershell_dependency_up == 0

# Worker capacity saturated
max without(instance) (mistershell_worker_slots_used)
  >= max without(instance) (mistershell_worker_slots_total)

Behaviour under partial failure

A scrape always answers with whatever can be gathered — it never fails outright because one subsystem is down:

  • If a backend dependency is unreachable, its mistershell_dependency_up reports 0 and the metric families that depend on it are absent from that scrape (treat absence as “unknown”, not zero).
  • The message-bus and stream families come from a periodic on-node sample. Right after a Core starts (or if sampling is interrupted) they are omitted until the next sample lands — usually within seconds.

Permissions

  • Scrape /metrics: app.fabric.read (via the bearer token’s user).