Skip to content
Stackship documentation Svenska

MonitoringAdministrators

Troubleshoot monitoring

Find out why charts are empty for every resource, why HTTP charts see no requests, why storage figures are missing or wrong, and why the monitoring tables keep growing.

Requires: monitoring/platform/read

The checks below use the Health page, which needs monitoring/platform/read at the root, and kubectl against the platform's own cluster. What each source gives is described in What monitoring needs from the cluster.

No resource has metrics

  1. On Health → Map, check that Monitoring, under Modules, is healthy, or run:

    bash
    kubectl -n stackship-system get deployment module-monitoring
  2. Read its log:

    bash
    kubectl -n stackship-system logs deployment/module-monitoring --since=15m
    Message Means
    Scrape failed for …; backing off A source could not be read; the message names it. It is tried again after at most 2 minutes
    Pod informer watch disconnected; reconnecting in 5s The module lost its watch on the cluster's pods. Until it reconnects, new pods are not filed under their resources
    Failed to attribute pod …; skipping One pod could not be filed under its resource; the others are unaffected
  3. If Health → Nodes shows CPU used and Memory used for the nodes, the kubelets are being read and the readings are not being filed under resources. Look for Failed to attribute pod in the log, and check that the resource's pods run in a namespace whose name starts with rg- and carry the labels platform.stackship.se/name and platform.stackship.se/boundary. A pod without platform.stackship.se/type is filed as an app, so any other resource needs that label too.

HTTP charts say no requests were observed

Every HTTP chart reading No requests observed in this period, on resources that do get traffic, means the module has not found the ingress controller's metrics.

bash
kubectl get endpoints -A | grep -E 'ingress-nginx-controller|traefik-metrics'
  • Nothing listed. On nginx, find the controller's Service and set IngressController__ServiceName to its name. On Traefik, add a Service named traefik-metrics for Traefik's metrics port 9100.
  • Listed without addresses. The Service selects no running pods.

Then restart the module — see Settings.

No storage on Health

The Storage card reads No storage provider reporting metrics on this cluster. when the module finds no Longhorn. That is expected on a cluster with another storage driver. With Longhorn, check that its Service exists and matches Longhorn__ServiceName:

bash
kubectl -n longhorn-system get endpoints longhorn-backend

A Monitoring target stale alert means the module finds Longhorn but gets no metrics from it.

Storage committed looks wrong

If Storage committed is far from what Longhorn itself reports — for example above 100 % on nodes where volumes still schedule — compare the module's over-provisioning percentage with Longhorn's, as shown in Longhorn. Set Longhorn__OverProvisioningPercentage to Longhorn's value and restart the module; the figure and the scheduling-headroom alert follow from then on.

The monitoring tables keep growing

Open Health → Database, which only a Platform Owner sees. A job listed under Background jobs that have never succeeded may be one of the module's retention, compression or roll-up jobs. Look in the module's log for Timescale policy statement did not apply, which names a policy it could not attach at start. Storage on the same tab shows how large each database has become. How long each table should keep its data is listed in Metric storage and retention.