Monitoring
What the platform measures about your resources, where you see it, and how alerts tell you that something needs attention.
The platform measures the resources it runs — CPU, memory, network traffic, disk use and HTTP requests — and keeps the history. Once a minute it checks those measurements and the state of each resource's pods against a set of built-in rules, and raises an alert when one of them finds a problem. There is nothing to set up: a resource is measured from the moment its pods run.
Where you see it
| Where | What |
|---|---|
| A resource's Overview, the Metrics card | Charts of the resource's usage and traffic, from the last 15 minutes up to the last 7 days — see Read a resource's metrics |
| A resource's Overview, the Active alerts card | The resource's alerts that are firing now. The card is only there while one is |
| A resource's Alerts tab | The resource's alerts, firing and resolved. The tab appears once the resource has raised one — see View alerts |
stsh monitor alerts |
The alerts firing now across a boundary |
| Health in the sidebar | For platform administrators: the platform's components, its nodes' capacity and the alerts about the platform itself — see The Health page |
What is measured
Every 15 seconds the platform reads the container statistics of every node, the fill level of every persistent volume and the request counters of the ingress controller, and files each reading under the resource it belongs to. What a resource gets depends on what it has:
- every resource with pods gets CPU, memory and network traffic;
- a resource with a persistent volume also gets disk use;
- a resource reached through the platform's ingress also gets HTTP requests per status class.
How the numbers are collected, and how long they are kept, is described in Metrics; which charts each resource type shows, in Charts and metrics.
What alerts cover
The built-in rules look for a container close to its CPU or memory limit or held back by its CPU limit, containers that crash, restart or are killed for running out of memory, images that cannot be pulled, pods that are not ready or cannot be scheduled, a rising share of HTTP server errors, and resources that have not been redeployed for 90 days. An alert resolves on its own once its rule no longer finds the problem. See Alerts and Alert rules.
What monitoring does not do
- It does not notify you. Alerts are shown in the portal and listed by the CLI and the API; the platform sends a resource's users no email, webhook or other message when one fires. The email digest it attempts for platform administrators is not delivered in this version — see Notifications.
- It has no rules of your own. The rules and their thresholds are built in and the same for every resource.
- It does not collect your application's own metrics or logs. It reads what the cluster reports about your containers, not an endpoint your application exposes. Logs are on each resource's own page — for an app, see Logs.
- It measures the platform's own cluster. A resource that runs in another cluster connected to the platform has no metrics and raises no alerts.
Security and operations analysis of your resources is a separate feature of the platform, Sentinel.
Pages
- Metrics — how the numbers are collected and kept
- Read a resource's metrics
- Charts and metrics
- Alerts — how an alert opens, changes and resolves
- View alerts
- Alert rules
- Permissions