What monitoring needs from the cluster
How the Monitoring module finds the kubelets, the ingress controller and Longhorn, what each gives, what is missing without them, and the settings that point it at them.
The Monitoring module runs as the deployment module-monitoring in stackship-system. It
watches the cluster it runs in — the platform's own — and nothing else: resources in other clusters
connected to the platform have no metrics and raise no alerts. It installs nothing into the
cluster; it finds what is there, and a component that is not there simply produces no data.
What it reads
| Source | Found as | What it gives | Without it |
|---|---|---|---|
| Each node's kubelet | Every node, through the Kubernetes API server's node proxy | Container CPU, memory and network; node CPU and memory in use; persistent volume fill levels | No metrics at all |
| nginx ingress controller | The pods behind a Service named ingress-nginx-controller, on port 10254 |
HTTP requests per resource and status | No HTTP charts and no ingress error-rate alerts |
| Traefik | The pods behind a Service named traefik-metrics, on port 9100 |
The same, from Traefik | The same |
| Longhorn | The pods behind a Service named longhorn-backend, on port 9500 |
Storage per node and disk, and volume health | No storage figures on Health and no storage alerts |
A Service is found by its name in any namespace. Node capacity — what each node can allocate and what its pods have requested — comes from the Kubernetes API itself.
The module reads all of this with its own service account, stackship-module-monitoring, which
the installer binds to the ClusterRole stackship-module-monitoring-role: read access to nodes,
pods, services, endpoints, persistent volumes, ingresses and Traefik IngressRoutes, and the node
proxy. The platform also binds it to a ClusterRole generated from the permission catalogue,
stackship-module-monitoring, which is read-only too. Neither grants write access.
Ingress controller
- nginx. The installer's nginx has its metrics turned on. If the controller was installed with
a Helm release name other than the chart's, its Service has another name — set
IngressController__ServiceNameto it. - Traefik. The installer's Traefik has the metrics Service
traefik-metricsturned on. A Traefik without that Service produces no HTTP metrics until a Service of that name exposes its metrics port 9100.
Requests are attributed through the labels the platform puts on its ingresses and Traefik routes.
Longhorn
Two Longhorn settings are not exported as metrics, so the module has to be told them, and they must match what Longhorn is set to. A mismatch skews the storage figures and alerts without any error.
| Setting | Must match Longhorn's | Default | Used for |
|---|---|---|---|
Longhorn__OverProvisioningPercentage |
storage-over-provisioning-percentage |
100 |
Storage committed on Health and the scheduling-headroom alert. With Longhorn at 200 and the module at 100, both read twice the real figure |
Longhorn__MinimalAvailablePercentage |
storage-minimal-available-percentage |
25 |
The disk-space alert, which warns 5 points above it |
Compare them with:
kubectl -n longhorn-system get settings.longhorn.io storage-over-provisioning-percentage -o jsonpath='{.value}'
kubectl -n longhorn-system get settings.longhorn.io storage-minimal-available-percentage -o jsonpath='{.value}'
kubectl -n stackship-system get configmap module-monitoring-config -o yaml | grep Longhorn__Settings
The installer asks for the first five among the Monitoring module's settings, under Ingress
and Storage. They reach the module as environment variables from the ConfigMap
module-monitoring-config in stackship-system, and are read when it starts.
| Environment variable | In the installer | Default | Meaning |
|---|---|---|---|
IngressController__ServiceName |
Ingress Controller Service Name | ingress-nginx-controller |
The nginx controller's Service |
Longhorn__ServiceName |
Longhorn Manager Service Name | longhorn-backend |
Longhorn's Service |
Longhorn__MetricsPort |
Longhorn Metrics Port | 9500 |
The port Longhorn serves its metrics on |
Longhorn__OverProvisioningPercentage |
Longhorn Storage Over-Provisioning % | 100 |
See Longhorn |
Longhorn__MinimalAvailablePercentage |
Longhorn Minimal Available % | 25 |
See Longhorn |
TraefikController__ServiceName |
Not asked | traefik-metrics |
Traefik's metrics Service |
TraefikController__MetricsPort |
Not asked | 9100 |
The port Traefik serves its metrics on |
Alerts__RetentionDays |
Not asked | 30 |
Days a resolved alert is kept — see Alert retention |
Alerts__Notifications__Enabled |
Not asked | true |
Whether the module tries to email Critical alerts — see Alert email |
After changing the ConfigMap, restart the module:
kubectl -n stackship-system rollout restart deployment/module-monitoringPlatform upgrades keep a value you have changed and a key you have added. Every setting the module reads is listed in its configuration reference.