Platform alerts
The alerts about the platform itself — storage, monitoring and conditions other platform components report — where they show, how they leave the cluster, and why no alert email arrives.
Platform alerts are about the platform rather than a tenant's resource. They belong to no boundary,
never appear in a tenant's alert lists, and are shown under Active platform alerts on the
Nodes tab of the Health page, which needs monitoring/platform/read at the
root. They open, change and resolve the same way as resource alerts — see
Alerts.
The list is also served by GET /admin/monitoring/alerts: the firing alerts, or with
includeResolved=true the resolved ones still kept as well, at most 200.
Storage
These rules read Longhorn's metrics and stay silent on a cluster without Longhorn. Their thresholds depend on two Longhorn settings the Monitoring module must be told — see Longhorn.
| Rule | Level | Fires when |
|---|---|---|
storage-scheduling-headroom — Storage scheduling headroom low |
Warning, Critical | A node's committed storage has stayed above 75 % of what Longhorn can schedule on it for the last 15 minutes; Critical above 90 % for the last 5 minutes. What it can schedule is the node's capacity minus reserved space, times the over-provisioning percentage. When it is used up, new volumes cannot be placed |
storage-disk-space — Storage disk nearly full |
Warning, Critical | A disk's free space, at its fullest over the last 15 minutes, is less than 5 percentage points above the minimal-available percentage; Critical when less than 1 point above. With the defaults: below 30 % and below 26 % free |
storage-node-ready — Storage node not ready, Storage disk not ready |
Critical | Longhorn has reported a node or a disk as not ready for the whole of the last 5 minutes. It contributes no storage |
storage-volume-health — Volume degraded, Volume faulted |
Warning, Critical | A volume has been degraded — rebuilding a replica — for the whole of the last 15 minutes, or faulted — no healthy replica, cannot be attached — for the last 2 minutes |
A volume alert names the resource group and the claim the volume backs when the platform can tell,
in its description and in metadata (resourceGroup, resourceName).
Monitoring itself
| Rule | Level | Fires when |
|---|---|---|
scrape-target-stale — Monitoring target stale |
Critical | Longhorn's metrics endpoints are known but no storage metrics have arrived for over 10 minutes. Every storage alert is blind while it fires |
Raised by other platform components
A platform component can raise and clear an alert of its own through the Monitoring module; it
needs monitoring/platform/alerts/write, which the service role Platform Alert Publisher holds.
Platform Owner and Platform Contributor hold it too, so they can raise and resolve such alerts
through the same API — see Permissions.
The platform's backups, for example, raise one while failing archiving of the platform database's
write-ahead log is filling its volume. These alerts have the category Platform and open and
resolve when the component says so.
Each time one of them opens, changes level or resolves, the Monitoring module also writes it to its
log as a record named stackship.alert, at the alert's level and never below Warning. With
telemetry export on, that record leaves the cluster with the platform's other logs, unless the
export's minimum log level is above it — see
Export platform telemetry. The rules on this page and the resource
rules are not written this way.
Alert email
Once a minute, after the rules have run, the Monitoring module tries to email one digest of the
alerts that opened at level Critical or resolved while at Critical. It covers its own rules, the
resource rules included, but not alerts raised by other components; an alert that opens at a lower
level and later rises to Critical is not announced when it rises. The digest goes to every user assigned
Platform Owner or Platform Contributor at the root directly — members of a group that holds the role
are not included — or, when there is no such user, to the addresses in
Alerts__Notifications__FallbackAddresses.
Important
No alert email is delivered in this version. The digest uses the email template
platform-alert-digest, which is not among the templates the platform ships, so the platform refuses every message and the module logsCould not enqueue an alert digestfor each recipient. Watch the Health page and the resources' Alerts instead.
Setting Alerts__Notifications__Enabled to false on the Monitoring module stops the attempts —
see Settings.