Skip to content
Stackship documentation Svenska

PostgreSQLAdministrators

What PostgreSQL backups need

How managed PostgreSQL clusters archive to the platform's backup target — the dependencies, the objects the operator projects into each namespace, the paths and retention — and what happens without a target.

A tenant's Enable Backups switch only works when the platform can write to a backup target. This page describes what that takes; the tenant side is in Backups and point-in-time recovery.

The dependencies

PostgreSQL backups use the CloudNativePG Barman Cloud plugin, which writes to the same object storage bucket Velero backs the platform up to. The installer installs both as platform dependencies: velero, and cnpg-barman-plugin, which depends on cert-manager, cnpg-operator and velero and is not optional.

The target itself is Velero's: the credentials Secret velero (key cloud, holding aws_access_key_id and aws_secret_access_key) and the BackupStorageLocation default, both in the namespace velero. From the location the operator takes the bucket, the region (us-east-1 when none is set), the endpoint URL and, when there is one, the endpoint's CA certificate. Change the target through the platform's backup settings — see Lifecycle — rather than by editing Velero's objects.

What the operator projects

Every two minutes the Stackship operator writes a credentials Secret and a Barman ObjectStore, both named stackship-backup-target, into:

  • the platform namespace, stackship-system, always;
  • every namespace that holds a CloudNativePG cluster archiving to stackship-backup-target — that is, every resource group with a PostgreSQL cluster whose backups are on.

It removes them again from a namespace once no such cluster is left in it. The ObjectStore points at:

Namespace Destination
The platform namespace s3://<bucket>/platform/postgres/
A resource group's namespace s3://<bucket>/boundaries/<boundary-id>/<namespace>/postgres/

WAL and base backups are compressed with gzip. Within that path each cluster archives under a name of its own: the cluster's name and a random six-character suffix, kept while backups stay on; turning backups off and on again starts a new name. Each restore moves the cluster to the next generation of that name — -g2, -g3 and so on — so that it never writes into an archive that already holds data.

Retention

Every projected ObjectStore carries the retention policy 30d, and the plugin applies it every 30 minutes: it keeps what is needed to recover to any moment of the last 30 days and removes older backups and WAL. This is the only retention that applies to a tenant's PostgreSQL backups; the Backup Retention Policy a tenant sets on a cluster, and the lifetime asked for when a snapshot is taken, do not change it.

The plugin applies it from the instances of a running cluster, to the archive that cluster writes to. Archives no cluster writes to any more — earlier generations after a restore, the old name after backups were turned off and on, and the archives of deleted clusters — are never pruned and stay in the bucket until someone removes them.

Scheduled snapshots

A cluster with backups on gets a BackupPolicy next to it, owned by the cluster. The operator renders it as a CloudNativePG ScheduledBackup named <cluster>-policy, with the policy's five-field cron in UTC. The first scheduled backup runs at the next scheduled time, not when the schedule is created, and pausing the policy suspends it. The scheduled backups are owned by the CloudNativePG cluster (backupOwnerReference: cluster), so a restore, which deletes and recreates the cluster, removes them; their data stays in the archive.

The platform's own backup

The platform's file-system backup with Velero skips the PostgreSQL volumes: every instance pod carries backup.velero.io/backup-volumes-excludes: pgdata,pg-wal. A PostgreSQL cluster is only ever recovered from its archive, which a copy of its volumes could not replace.

Without a target

When Velero's Secret or storage location is missing, or names no bucket or no credentials, the operator projects nothing — and removes nothing it projected before. A tenant can still turn backups on, but the cluster has nowhere to archive to: no snapshot completes, a snapshot or a restore is refused once the cluster reports archiving as failing, and point-in-time recovery has no window.

While archiving fails, PostgreSQL keeps every WAL segment it has not archived, so the cluster's volumes fill up until the target works again. The same holds when a working target becomes full or unreachable.