Skip to content
Stackship documentation Svenska

MonitoringUsers

Alerts

How the platform decides to raise an alert on a resource, how an open alert changes and resolves, what its level means, and how long alerts are kept.

An alert says that a built-in rule has found a problem with one of your resources: a container near its memory limit, a pod that cannot be scheduled, a route answering with server errors. Alerts are raised and resolved by the platform alone; there is nothing to configure and nothing to close.

How an alert is raised

Once a minute the platform runs every rule against the last few minutes of metrics and the current state of the resources' pods. A rule that finds a problem opens an alert on the resource. The rules and their thresholds are listed in Alert rules.

An alert is about one finding, at the granularity its rule judges: one container of the resource, the resource's pods as a whole, or one route of its HTTP traffic. Two containers of one resource that both run out of memory give two alerts.

While it is open

As long as the rule keeps finding the problem, the alert stays open and keeps the time it was triggered. Its level, description and details follow the latest check: a container that goes from 80 % to 96 % of its limit turns the same alert from Warning to Critical rather than opening a second one.

How it resolves

The first minute the rule no longer finds the problem, the alert is resolved and gets the time it resolved. Alerts cannot be acknowledged, silenced or closed by hand. If the problem comes back later, a new alert opens.

Levels

Level Raised for Badge
Critical Something needs action now: 95 % of a limit, a quarter of requests failing Red
Error Something is failing: a crash loop, an out-of-memory kill, an image that cannot be pulled, 90 % of a limit Red
Warning Something is degraded or heading for trouble: 70 % of a limit, repeated restarts, pods not ready or not scheduled Amber
Information, Recommendation, Insight Levels the platform knows; no built-in rule raises them Grey

How long alerts are kept

A resolved alert is kept for 30 days after it resolved, unless your platform is configured otherwise, and then deleted. An alert that is still firing is kept however long it lasts. A resource's Alerts tab shows at most 50 alerts: the firing ones first, then the most recent.

Notifications

The platform does not tell a resource's users about an alert: it sends them no email, webhook or chat message. Alerts are on the resource's page, and stsh monitor alerts and the API list them for your own tools to poll — see View alerts.

The only message the platform attempts is an email digest of Critical alerts, resource alerts included, to its Platform Owners and Platform Contributors. In this version that email is not delivered — see Alert email.

Alerts about the platform

Some alerts are about the platform rather than any resource of yours: its storage, its own backups and monitoring itself. They belong to no boundary, never appear on a resource's page, and are shown to platform administrators on the Health page — see Platform alerts.

Pages in this section