Skip to content
Stackship documentation Svenska

MonitoringAdministrators

The Health page

What the Health page's Map, Nodes and Database tabs show about the platform's components, its nodes' capacity and its own database, and who can open each.

Requires: monitoring/platform/read

Health shows the platform itself rather than any tenant's resources: what it is made of and which piece is broken, how much capacity its nodes have left, and how its own database is doing. Open it from Health in the sidebar or the Health card on the Admin page.

Who sees what

All of these are checked at the platform root.

Part Needs Held by
The page, and the Nodes tab monitoring/platform/read Platform Reader, Platform Contributor, Platform Owner
The Map tab kernel/platform/read Platform Reader, Platform Contributor, Platform Owner
The Database tab, and a component's pods and events on Map kernel/platform/diagnose Platform Owner

The Database tab is not shown to those without kernel/platform/diagnose. Every tab refreshes every 30 seconds; the 24-hour sparklines on Nodes every 5 minutes.

Map

The platform's components, grouped in tiers from the outside in — Ingress, Edge, Modules, Reconcile and Data — for each cluster connected to the platform.

  • A banner sums up the state: All systems healthy, or how many components need attention, with the counts of healthy components, components with issues and components not in use.
  • Each component shows its status — healthy, degraded, down, not installed, not in use or unknown — how many of its replicas are ready, and its version when the Lifecycle Manager can say. A component that is not healthy gives the reason.
  • Hide components not in use leaves out the components marked not in use.
  • The platform's own cluster is marked This cluster. Components in other clusters are not observed from here: their frame says so, and for an agent cluster shows when its agent was last seen, or that its agent identity has been revoked.

A Platform Owner can select a component to see its pods — phase, readiness, restarts, node, and what stops a pod from running — its recent events, and why it is stuck when it is. Component details refresh every 10 seconds.

Nodes

Capacity of the platform's own cluster, from the Monitoring module.

Part Shows
CPU, Memory and Storage cards The whole cluster: Committed and Used against the total. Storage has no used figure; its Used bar shows a dash
Nodes table Per node: CPU committed, CPU used, the committed CPU over the last 24h, Memory committed, Memory used, the committed memory over the last 24h, and Storage committed. The node with the most committed dimension comes first; a figure turns amber at 75 % and red at 90 %
Unhealthy volumes Storage volumes that are degraded or faulted, with the resource group and claim each belongs to, or unattributed. Only shown when there are any
Active platform alerts The platform alerts firing now — see Platform alerts. Only shown when there are any
  • Committed is what the pods placed on a node have reserved — the sum of their CPU and memory requests, including system pods — against what the node can allocate to pods. It decides whether the next pod fits.
  • Used is what the node measures in use, against the same total. It can be far below committed.
  • Storage committed is the storage Longhorn has promised to volume replicas on the node, against the node's storage capacity minus its reserved space, multiplied by the Longhorn over-provisioning percentage the Monitoring module is set to. When that setting does not match Longhorn's own, this figure is off by the same ratio — see Longhorn.
  • Without Longhorn the Storage column is left out and its card reads No storage provider reporting metrics on this cluster.

Database

The platform's own PostgreSQL cluster, which holds every module's data. Each part shows what it could read; Not everything could be read lists what it could not, and why.

Part Shows
Background jobs that have never succeeded At the top, in red: database background jobs that have run and never once succeeded. The Monitoring module's retention, compression and roll-ups run as such jobs — see Metric storage and retention
A failing archive is filling this database's volume Shown while the platform's backups report that failing write-ahead-log archiving is filling the volume, with the link Suspend or resume archiving on Platform → Backups
The cluster Name, phase and image; Failover in progress from one instance to another; each instance as primary or replica and whether it is healthy; Last successful backup; Recoverable from; the link Open Platform → Backups
Connections Connections In use against the maximum, the count per state, Longest transaction and Waiting on locks. A note says when query text is hidden because the platform's database role cannot read other roles' queries
Storage Database size against volume capacity, and the size of each database. The size leaves out write-ahead log and temporary files, so the volume holds more than it shows
Background jobs Every background job, with how many of its runs succeeded