Skip to content
Stackship documentation Svenska

PostgreSQLUsers

Find out why a cluster is not ready

Follow a PostgreSQL cluster's operation instance by instance, check its health, and act on the reasons the platform gives for an instance that does not start.

Requires: postgrescluster/read

Follow the operation

Creating, starting, restarting and changing a cluster each start an operation. Open the cluster's Operations tab: the operation has a step per instance — Create instance 1 of 3, Start instance 2 of 3 — and a last step that waits for the whole cluster, such as Wait for the cluster to report healthy. A restart has a single step, Restart 3 instance(s) (rolling). See Operations.

When an instance does not come up, the open step says why — see the table below. If the problem lasts, the operation fails with that reason instead of waiting on.

Check the health

bash
stsh pg health orders-db

The answer is healthy: true, or false with a reason: the cluster is stopped, how many instances are ready, the operator's own status, or the same reason the operation would give. Through the API: GET .../postgresclusters/<name>/health.

What the reasons mean

The platform says What to do
Pod … cannot be scheduled: … No node has the CPU or memory left for the instance. Choose a smaller plan, or ask the platform's operators for capacity.
Pod … cannot pull its image … The image named by the version cannot be fetched. Check the version you set through the API, or ask the operators whether the cluster can reach the image registry.
Volume … is not bound … The storage class has not created the volume. The operators need to look at the storage provider.
Volume … has not grown from … to … The storage provider cannot grow the volume, usually for lack of room. Until it does, the cluster applies no other change of any kind. The operators need to make room on the storage provider.
Pod … keeps running out of memory (OOMKilled). Raise the memory limit. Move to a larger plan, or raise Memory Limit — see Tune resources individually.
Pod … container … keeps exiting (CrashLoopBackOff …) PostgreSQL does not start. A parameter you set is the usual cause; undo it — see Set PostgreSQL parameters.
Pod … container … cannot start (CreateContainerConfigError …) The instance's container cannot be set up, usually because a Secret or setting it refers to is missing. Ask the platform's operators.
Pod … failed (…) The instance's pod stopped for good. The reason in brackets comes from Kubernetes; ask the platform's operators if it does not explain itself.

When you change or restart a cluster, the platform recreates an instance that is stuck as it is — on an image it cannot pull, in a crash loop, or with a container it cannot set up — so that the new configuration applies to it. The cluster's Activity then records Recreated instance(s) … so the current configuration applies.

Changes that were not applied in full

Some changes are applied in part: a smaller storage size, a different storage class, a new database name, removing the WAL volume, or more instances or storage than the provider has room for. The platform keeps the current value for that part, applies the rest, and says so — the API and the CLI return it in warnings.

Logs

The portal and the CLI do not show a database's logs. When the operation, the health check and the Activity do not explain a problem, the platform's operators can read the instances' logs in the Kubernetes cluster.