Troubleshoot monitoring
Find out why charts are empty for every resource, why HTTP charts see no requests, why storage figures are missing or wrong, and why the monitoring tables keep growing.
Requires: monitoring/platform/read
The checks below use the Health page, which needs monitoring/platform/read at the root, and
kubectl against the platform's own cluster. What each source gives is described in
What monitoring needs from the cluster.
No resource has metrics
On Health → Map, check that Monitoring, under Modules, is healthy, or run:
kubectl -n stackship-system get deployment module-monitoringRead its log:
kubectl -n stackship-system logs deployment/module-monitoring --since=15mMessage Means Scrape failed for …; backing offA source could not be read; the message names it. It is tried again after at most 2 minutes Pod informer watch disconnected; reconnecting in 5sThe module lost its watch on the cluster's pods. Until it reconnects, new pods are not filed under their resources Failed to attribute pod …; skippingOne pod could not be filed under its resource; the others are unaffected If Health → Nodes shows CPU used and Memory used for the nodes, the kubelets are being read and the readings are not being filed under resources. Look for
Failed to attribute podin the log, and check that the resource's pods run in a namespace whose name starts withrg-and carry the labelsplatform.stackship.se/nameandplatform.stackship.se/boundary. A pod withoutplatform.stackship.se/typeis filed as an app, so any other resource needs that label too.
HTTP charts say no requests were observed
Every HTTP chart reading No requests observed in this period, on resources that do get traffic, means the module has not found the ingress controller's metrics.
kubectl get endpoints -A | grep -E 'ingress-nginx-controller|traefik-metrics'- Nothing listed. On nginx, find the controller's Service and set
IngressController__ServiceNameto its name. On Traefik, add a Service namedtraefik-metricsfor Traefik's metrics port 9100. - Listed without addresses. The Service selects no running pods.
Then restart the module — see Settings.
No storage on Health
The Storage card reads No storage provider reporting metrics on this cluster. when the
module finds no Longhorn. That is expected on a cluster with another storage driver. With Longhorn,
check that its Service exists and matches Longhorn__ServiceName:
kubectl -n longhorn-system get endpoints longhorn-backendA Monitoring target stale alert means the module finds Longhorn but gets no metrics from it.
Storage committed looks wrong
If Storage committed is far from what Longhorn itself reports — for example above 100 % on
nodes where volumes still schedule — compare the module's over-provisioning percentage with
Longhorn's, as shown in Longhorn. Set
Longhorn__OverProvisioningPercentage to Longhorn's value and restart the module; the figure and
the scheduling-headroom alert follow from then on.
The monitoring tables keep growing
Open Health → Database, which only a Platform Owner sees. A job listed under Background jobs
that have never succeeded may be one of the module's retention, compression or roll-up jobs. Look in the module's log for
Timescale policy statement did not apply, which names a policy it could not attach at start.
Storage on the same tab shows how large each database has become. How long each table should
keep its data is listed in Metric storage and retention.