What PostgreSQL backups need
How managed PostgreSQL clusters archive to the platform's backup target — the dependencies, the objects the operator projects into each namespace, the paths and retention — and what happens without a target.
A tenant's Enable Backups switch only works when the platform can write to a backup target. This page describes what that takes; the tenant side is in Backups and point-in-time recovery.
The dependencies
PostgreSQL backups use the CloudNativePG Barman Cloud plugin, which writes to the same object
storage bucket Velero backs the platform up to. The installer installs both as platform
dependencies: velero, and cnpg-barman-plugin, which depends on cert-manager, cnpg-operator
and velero and is not optional.
The target itself is Velero's: the credentials Secret velero (key cloud, holding
aws_access_key_id and aws_secret_access_key) and the BackupStorageLocation default, both in
the namespace velero. From the location the operator takes the bucket, the region (us-east-1
when none is set), the endpoint URL and, when there is one, the endpoint's CA certificate. Change
the target through the platform's backup settings — see Lifecycle — rather
than by editing Velero's objects.
What the operator projects
Every two minutes the Stackship operator writes a credentials Secret and a Barman ObjectStore,
both named stackship-backup-target, into:
- the platform namespace,
stackship-system, always; - every namespace that holds a CloudNativePG cluster archiving to
stackship-backup-target— that is, every resource group with a PostgreSQL cluster whose backups are on.
It removes them again from a namespace once no such cluster is left in it. The ObjectStore points
at:
| Namespace | Destination |
|---|---|
| The platform namespace | s3://<bucket>/platform/postgres/ |
| A resource group's namespace | s3://<bucket>/boundaries/<boundary-id>/<namespace>/postgres/ |
WAL and base backups are compressed with gzip. Within that path each cluster archives under a name
of its own: the cluster's name and a random six-character suffix, kept while backups stay on;
turning backups off and on again starts a new name.
Each restore moves the cluster to the next generation of that name — -g2, -g3 and so on — so
that it never writes into an archive that already holds data.
Retention
Every projected ObjectStore carries the retention policy 30d, and the plugin applies it every
30 minutes: it keeps what is needed to recover to any moment of the last 30 days and removes older
backups and WAL. This is the only retention that applies to a tenant's PostgreSQL backups; the
Backup Retention Policy a tenant sets on a cluster, and the lifetime asked for when a snapshot is
taken, do not change it.
The plugin applies it from the instances of a running cluster, to the archive that cluster writes to. Archives no cluster writes to any more — earlier generations after a restore, the old name after backups were turned off and on, and the archives of deleted clusters — are never pruned and stay in the bucket until someone removes them.
Scheduled snapshots
A cluster with backups on gets a BackupPolicy next to it, owned by the cluster. The operator
renders it as a CloudNativePG ScheduledBackup named <cluster>-policy, with the policy's
five-field cron in UTC. The first scheduled backup runs at the next scheduled time, not when the
schedule is created, and pausing the policy suspends it. The scheduled backups are owned by the
CloudNativePG cluster (backupOwnerReference: cluster), so a restore, which deletes and recreates
the cluster, removes them; their data stays in the archive.
The platform's own backup
The platform's file-system backup with Velero skips the PostgreSQL volumes: every instance pod
carries backup.velero.io/backup-volumes-excludes: pgdata,pg-wal. A PostgreSQL cluster is only ever
recovered from its archive, which a copy of its volumes could not replace.
Without a target
When Velero's Secret or storage location is missing, or names no bucket or no credentials, the operator projects nothing — and removes nothing it projected before. A tenant can still turn backups on, but the cluster has nowhere to archive to: no snapshot completes, a snapshot or a restore is refused once the cluster reports archiving as failing, and point-in-time recovery has no window.
While archiving fails, PostgreSQL keeps every WAL segment it has not archived, so the cluster's volumes fill up until the target works again. The same holds when a working target becomes full or unreachable.