Backups and point-in-time recovery
How PostgreSQL backups work — the continuous archive, snapshots, the recovery window, retention — and what a restore does to the cluster.
The archive and snapshots
With backups on, a cluster keeps an archive in the platform's backup store, the object storage the platform's operators set up for backups. Two things go into it:
- the write-ahead log (WAL), shipped continuously as the database writes it. A log segment is archived when it is full, and by default after five minutes at the latest, so the archive trails the database by a few minutes at most;
- base backups — full copies of the database, taken while it runs. These are the cluster's snapshots. One is taken on the backup schedule, and you can take one at any time.
The snapshots are listed on the cluster's Snapshots tab, marked Database backup. A snapshot is consistent — the database as it was at one moment — and nothing in the cluster has to stop for it. Copies of the volumes are not used for PostgreSQL.
Two ways to restore
- Restore a snapshot — the database comes back as it was when the snapshot was taken.
- Restore to a point in time — the database comes back as it was at a moment you choose: the platform starts from a snapshot taken before that moment and replays the archived log up to it.
The recovery window is the stretch of time a point-in-time restore can reach: from the oldest point the archive can recover to, until now. The Snapshots tab shows where it starts — "Snapshots plus a continuous archive since …" — once the archive holds a snapshot.
What a restore does
A restore replaces the cluster's data in place:
- every instance is shut down;
- the cluster is removed and created again under the same name from the archive, at the snapshot or the chosen moment;
- its instances start and the cluster is Running again.
The connection string, the service names and the password stay the same, so applications need no change; they only have to reconnect. With external access on, the load balancer is created again and its address — and with it the external connection string — can change. Expect a few minutes without the database, more for a large one, and everything written after the restored moment is gone. The cluster's Operations tab follows the restore.
If recovery from the archive fails, or the restore has not finished 45 minutes after it started, the platform recreates the cluster from the newest snapshot in the archive it was writing before the restore — not at the moment the restore started, so what was written after that snapshot is missing — and the operation reports the failure.
Caution
Other failures after the old cluster has been removed — for example a snapshot that can no longer be found — end the restore with no cluster at all, and so does a fallback that fails in turn. The cluster's volumes are gone by then; only the archive holds the data, and the platform's operators have to recover it.
After a restore, the cluster writes to a new archive, and the recovery window starts again with the cluster's next snapshot. A point-in-time restore cannot reach back into the archive used before.
Caution
A restore removes the scheduled snapshots — those described as Scheduled by … — together with the old cluster. Snapshots taken by hand stay in the list and can still be restored as they are. Restoring a scheduled snapshot runs into the same removal and can end with no cluster; take a snapshot by hand and restore that one where you can.
Retention
Important
Today every cluster's archive keeps what it needs to recover to any moment of the last 30 days, and removes what is older. This is set on the backup store, not per cluster: the cluster's Backup Retention Policy is stored on the cluster but not applied, and the Keep for you choose when you take a snapshot has no effect.
Deleting a snapshot removes it from the list; its data stays in the archive until retention removes it. Retention is applied by the running cluster to the archive it writes to, so archives that no cluster writes to any more — the old archive after a restore or after backups were turned off and on again, and the archive of a deleted cluster — are not pruned and stay in the backup store.
When archiving fails
When the archive cannot be written to — the backup store is full or unreachable, or the platform has no backup target — no snapshot completes, and once the cluster reports archiving as failing, taking a snapshot and restoring are refused. PostgreSQL also keeps every log segment it has not archived yet, so the cluster's volumes fill up for as long as archiving fails. The platform's operators are the ones to fix the backup store.
Replicas are not backups
Replicas protect against a failed instance, not against a mistake: a table dropped on the primary is dropped on every replica a moment later. Only the archive can bring it back.