Instances and high availability
How a PostgreSQL cluster's primary and replicas work, which service leads where, what happens on a failover, a restart and a stop, and what it means for connections.
Primary and replicas
Every instance of a cluster holds a full copy of the database on its own volumes. One of them is the primary: it takes all writes. The others are replicas: they stream the primary's write-ahead log and replay it, and they can answer reads. The Overview shows which instance is the primary under Current Primary, and how many instances are ready under Ready Instances.
Replication is asynchronous. The primary confirms a commit without waiting for the replicas, so a replica is usually a moment behind, and a transaction committed just before the primary fails may not have reached any replica.
The plans nano, small and medium run one instance; standard and larger run three. Any count
from 1 to 10 can be set — see Change the number of instances.
Services
The operator keeps three addresses in the cluster's resource group, all on port 5432:
| Service | Leads to | Shown on the Overview as |
|---|---|---|
<cluster>-rw |
The primary | Write Service |
<cluster>-r |
Any instance, the primary included | Read Service |
<cluster>-ro |
The replicas only | — |
The connection string uses <cluster>-rw, so it follows the primary wherever it moves. With
PgBouncer on, <cluster>-pooler-rw leads to the pooler, which forwards to the primary.
Failover
When the primary fails — its node goes down, its pod dies, it stops answering — the operator
promotes the replica that is furthest ahead and moves <cluster>-rw to it. Connections to the old
primary are cut; an application that reconnects reaches the new one. Transactions the old primary
had not yet sent to that replica are lost.
A cluster with one instance has nothing to fail over to: until the instance is running again, elsewhere if its node is gone, the database is unavailable.
The platform offers no action to move the primary on purpose.
Planned restarts
A Restart, a change of CPU, memory or instance count, and a change to a parameter that needs a restart all restart the instances one at a time: the replicas first, then the primary. A cluster with more than one instance keeps serving throughout, and connections to the primary drop once while it restarts. A single instance is briefly unavailable. The portal says which of the two a plan change will be before you confirm it.
A new major version is different: it stops every instance, the primary included, until the upgrade has run.
Stopped clusters
Stop shuts down every instance and keeps the volumes. A stopped cluster accepts no connections, its Data Explorer is unavailable, and its status reads Stopped until Start brings it back.