Skip to content
Stackship documentation Svenska

MonitoringUsers

Alert rules

Every built-in rule that raises an alert on a resource, the level it raises, when it fires and when it resolves, and the details each alert carries.

Every rule runs once a minute over the pods of your resources in the platform's own cluster. The rules and their thresholds are the same for every resource and cannot be changed. The rule id is the alert's alertId in the API; the name is what the portal shows.

Resource usage

Rule Level Fires when Resolves when
resource-exhaustion — Resource exhaustion Warning at 70 % and 80 %, Error at 90 %, Critical at 95 % A container's CPU or memory use reaches that share of its limit. CPU is the latest per-second use in the last 5 minutes, memory the latest working-set reading — each the most recent reading of that container from any of its replicas; the higher of the two decides Both are below 70 %
cpu-throttling — CPU throttling Warning above 25 %, Error above 50 % A container was held back by its CPU limit in more than that share of its CPU scheduling periods over the last 10 minutes, counted over all its replicas The share is 25 % or less
container-oom-killed — Container out of memory Error A container was killed for going over its memory limit in the last 30 minutes 30 minutes after the latest kill

Resource exhaustion and CPU throttling only judge containers that declare both a CPU and a memory limit. A throttled container can use far less than its limit on average, which is why throttling is a rule of its own. A resource with several replicas can show the same Resource exhaustion alert several times for one container — once for each replica running when it opened — with the same figures; the copies change and resolve together.

Containers and pods

Rule Level Fires when Resolves when
container-crash-loop — Container crash looping Error A container is in CrashLoopBackOff, or has restarted at least 3 times and its last run failed — a non-zero exit code or out of memory — within the last 10 minutes No replica of the container has failed in the last 10 minutes
container-restarting — Container restarting Warning A container has restarted at least 3 times and last restarted within the last 30 minutes, without being in a crash loop 30 minutes pass without a restart
image-pull-failing — Image pull failing Error A container cannot start because its image cannot be pulled: ImagePullBackOff, ErrImagePull or InvalidImageName. Fires at the first check The image is pulled
pod-not-ready — Pod not ready Warning Running pods of the resource have failed their readiness check for more than 10 minutes, so traffic is not reaching them. Pods with a crash-looping container are left to that rule No running pod has been unready for over 10 minutes: they are ready again, crash-looping or no longer running
pod-unschedulable — Pod cannot be scheduled Warning The cluster has found no node for pods of the resource for more than 5 minutes. The description carries the scheduler's message, cut to 512 characters — too little capacity, no matching node, a taint The pods are placed

A crash-looping container does not also raise Container restarting or Pod not ready; when it stops crashing but has restarted recently, Container restarting may follow at the lower level. A container that crash-loops because it keeps running out of memory raises both Container crash looping and Container out of memory. The container rules give one alert per container of the resource; the pod rules one per resource.

HTTP traffic

Rule Level Fires when Resolves when
ingress-error-rate — Ingress error rate Warning at 1 %, Error at 5 %, Critical at 25 % That share of the requests to the resource over the last 5 minutes got a 5xx status, and it had more than a trickle of traffic — see below The share is below 1 %, or traffic falls below that floor

The rule reads the platform's ingress controller. On nginx it judges each host and path of the resource separately; on Traefik the resource as a whole. It skips a route whose request rates, added up over every point in the 5 minutes, come to less than 1. That floor is well below one request per second, and the totalRps detail is that sum rather than a rate.

Lifecycle

Rule Level Fires when Resolves when
resource-not-updated-90d — Resource not updated in 90 days Warning In the pod of the resource the rule judges, the newest start of a container's previous run is more than 90 days ago That start is 90 days old or less, for example after a redeploy

The rule gives one alert per resource and judges one of its pods. It reads when the containers' runs before the current ones started, so it only sees pods where a container has restarted at least once. A resource whose containers have run without a restart since they were deployed never raises it, however old they are. It only judges resources with at least one container that declares both a CPU and a memory limit.

Details each alert carries

The API returns these in the alert's metadata:

Rule Details
resource-exhaustion container, cpuPercent, memoryPercent
cpu-throttling container, throttledPercent, cpuLimitCores
container-oom-killed container, exitCode, restartCount, finishedAt, affectedPods
container-crash-loop container, restartCount, lastTerminationReason, lastExitCode, affectedPods
container-restarting container, restartCount, lastTerminationReason, lastExitCode, lastRestartAt, affectedPods
image-pull-failing container, reason, message, affectedPods
pod-not-ready notReadyPods, pods, notReadySince
pod-unschedulable unschedulablePods, pods, reason, message, unschedulableSince
ingress-error-rate errorRatePercent, totalRps, and on nginx host and path
resource-not-updated-90d ageDays, lastUpdatedAt

pods lists at most ten pod names, followed by +N more when there are more. Alerts about the platform itself — storage and monitoring — are listed in Platform alerts.