Skip to content
Stackship documentation Svenska

MonitoringUsers

Metrics

Where a resource's metrics come from, how a reading is filed under its resource, what the percentages are measured against, and how long the history is kept.

A resource's metrics are readings the platform takes itself, from what the cluster reports about the resource's containers, volumes and traffic. Your application does not have to expose anything.

Where the numbers come from

Every 15 seconds the platform reads:

Source What it gives
The container statistics of each node CPU time, memory and network traffic of every container
The volume statistics of each node Used and total space of every persistent volume that a pod has mounted
The ingress controller, nginx or Traefik How many requests each route answered, per HTTP status

It keeps a short, fixed list of series from these sources and drops everything else — see Stored series.

How a reading finds its resource

Every pod the platform starts for a resource carries labels that name the resource, its type and its boundary, and so do the ingresses and Services it creates for one. A container reading is filed under the resource its pod names, a volume reading under the resource whose pod mounts the volume, and a request under the resource its ingress names — on Traefik, the resource of the Service behind the route. A container or request reading that cannot be placed that way is dropped. Two consequences:

  • An app's slots count as the app. A slot's pods carry the app's name, so the app's charts include its slots.
  • Only the platform's own cluster is read. A resource running in another cluster connected to the platform has no metrics.

Percent of what

CPU and memory are shown as a share of the resource's limits: the ceiling the cluster enforces by throttling CPU and by killing a container that exceeds its memory. 100 % is the wall the workload hits.

  • The limits are those of all the resource's containers that declare both a CPU and a memory limit. Each point is divided by the average limit of one of today's pods times the number of pods that reported at that point, so a scale-out does not distort the curve.
  • The usage counts every container of those pods, including one that declares no limits, such as a sidecar; only the limits leave it out. A pod with such a container reads higher than its limited containers alone would.
  • The limits used are the current ones. After a change of compute plan, earlier points are also measured against the new limits.
  • CPU can read above 100 %: the limit is enforced over short periods, and an average across a point can exceed it.
  • A resource none of whose containers declares both limits is shown in cores and bytes instead.

Disk is a share of the volume's capacity at each point, so a volume that has been enlarged reads correctly before and after the change.

Resolution and history

Readings are stored as taken and rolled up into 1-minute, 5-minute, hourly and daily summaries. The length of the range you ask for decides which one is read, and each is kept for a fixed time:

Range One point per Available for the last
Up to 1 hour 30 seconds 7 days
Up to 1 day 1 minute 7 days
Up to 7 days 5 minutes 30 days
Up to 90 days 1 hour 180 days
Longer 1 day 2 years

The portal's ranges are 15 minutes, 1 hour, 6 hours, 1 day and 7 days. Longer ranges, and ranges in the past, are read through the API — see Read a resource's metrics. The resolution follows the length of the range, not its age: a one-hour range that ended more than 7 days ago returns nothing.