Skip to content
Stackship documentation Svenska

MonitoringAdministrators

What monitoring needs from the cluster

How the Monitoring module finds the kubelets, the ingress controller and Longhorn, what each gives, what is missing without them, and the settings that point it at them.

The Monitoring module runs as the deployment module-monitoring in stackship-system. It watches the cluster it runs in — the platform's own — and nothing else: resources in other clusters connected to the platform have no metrics and raise no alerts. It installs nothing into the cluster; it finds what is there, and a component that is not there simply produces no data.

What it reads

Source Found as What it gives Without it
Each node's kubelet Every node, through the Kubernetes API server's node proxy Container CPU, memory and network; node CPU and memory in use; persistent volume fill levels No metrics at all
nginx ingress controller The pods behind a Service named ingress-nginx-controller, on port 10254 HTTP requests per resource and status No HTTP charts and no ingress error-rate alerts
Traefik The pods behind a Service named traefik-metrics, on port 9100 The same, from Traefik The same
Longhorn The pods behind a Service named longhorn-backend, on port 9500 Storage per node and disk, and volume health No storage figures on Health and no storage alerts

A Service is found by its name in any namespace. Node capacity — what each node can allocate and what its pods have requested — comes from the Kubernetes API itself.

The module reads all of this with its own service account, stackship-module-monitoring, which the installer binds to the ClusterRole stackship-module-monitoring-role: read access to nodes, pods, services, endpoints, persistent volumes, ingresses and Traefik IngressRoutes, and the node proxy. The platform also binds it to a ClusterRole generated from the permission catalogue, stackship-module-monitoring, which is read-only too. Neither grants write access.

Ingress controller

  • nginx. The installer's nginx has its metrics turned on. If the controller was installed with a Helm release name other than the chart's, its Service has another name — set IngressController__ServiceName to it.
  • Traefik. The installer's Traefik has the metrics Service traefik-metrics turned on. A Traefik without that Service produces no HTTP metrics until a Service of that name exposes its metrics port 9100.

Requests are attributed through the labels the platform puts on its ingresses and Traefik routes.

Longhorn

Two Longhorn settings are not exported as metrics, so the module has to be told them, and they must match what Longhorn is set to. A mismatch skews the storage figures and alerts without any error.

Setting Must match Longhorn's Default Used for
Longhorn__OverProvisioningPercentage storage-over-provisioning-percentage 100 Storage committed on Health and the scheduling-headroom alert. With Longhorn at 200 and the module at 100, both read twice the real figure
Longhorn__MinimalAvailablePercentage storage-minimal-available-percentage 25 The disk-space alert, which warns 5 points above it

Compare them with:

bash
kubectl -n longhorn-system get settings.longhorn.io storage-over-provisioning-percentage -o jsonpath='{.value}'
kubectl -n longhorn-system get settings.longhorn.io storage-minimal-available-percentage -o jsonpath='{.value}'
kubectl -n stackship-system get configmap module-monitoring-config -o yaml | grep Longhorn__

Settings

The installer asks for the first five among the Monitoring module's settings, under Ingress and Storage. They reach the module as environment variables from the ConfigMap module-monitoring-config in stackship-system, and are read when it starts.

Environment variable In the installer Default Meaning
IngressController__ServiceName Ingress Controller Service Name ingress-nginx-controller The nginx controller's Service
Longhorn__ServiceName Longhorn Manager Service Name longhorn-backend Longhorn's Service
Longhorn__MetricsPort Longhorn Metrics Port 9500 The port Longhorn serves its metrics on
Longhorn__OverProvisioningPercentage Longhorn Storage Over-Provisioning % 100 See Longhorn
Longhorn__MinimalAvailablePercentage Longhorn Minimal Available % 25 See Longhorn
TraefikController__ServiceName Not asked traefik-metrics Traefik's metrics Service
TraefikController__MetricsPort Not asked 9100 The port Traefik serves its metrics on
Alerts__RetentionDays Not asked 30 Days a resolved alert is kept — see Alert retention
Alerts__Notifications__Enabled Not asked true Whether the module tries to email Critical alerts — see Alert email

After changing the ConfigMap, restart the module:

bash
kubectl -n stackship-system rollout restart deployment/module-monitoring

Platform upgrades keep a value you have changed and a key you have added. Every setting the module reads is listed in its configuration reference.