Manage platform backups
Take a platform backup now, change where backups go, change or pause the daily schedule, and keep a database from filling up behind a failing archive.
Requires: lifecycle/execute, lifecycle/backupsManage
Everything here is under Settings → Backups. Back up now needs lifecycle/execute at the
platform root, which Platform Owner and Platform Contributor hold. Changing the target or the
schedule, and suspending or resuming archiving, needs lifecycle/backupsManage, which of the
platform roles only Platform Owner holds: a wrong target loses the way back for the whole platform.
Take a backup now
Choose Back up now. The platform starts a backup named platform-manual-<date>-<time>, with the
same scope as the daily backup, and it appears in the list.
Velero runs backups one at a time, so a new one is refused while a backup taken with
Back up now is still queued or running; one started more than three hours ago no longer
counts. A daily backup that is running does not refuse it — the new one queues behind it. It is
refused as well when the platform-daily schedule is missing, because a manual backup takes its
scope from it.
Change the backup target
- On the Target card, choose Edit target.
- Fill in Bucket, Region, Endpoint, Checksum algorithm, Access key id and
Secret access key. The form starts from the current target, except the secret key, which is
never shown again after it is saved.
- Endpoint — leave it empty for AWS S3; for MinIO, Ceph or another S3-compatible store, the
full
https://URL. - Region — S3-compatible stores usually accept any region.
- Checksum algorithm — leave it empty for MinIO releases older than mid-2024, SeaweedFS and Ceph RGW, which reject the default checksums.
- Endpoint — leave it empty for AWS S3; for MinIO, Ceph or another S3-compatible store, the
full
- Choose Test connection. The store is asked to list the bucket with these credentials, and its answer is shown. Listing is all it checks: a key that may read the bucket but not write to it passes, and the backups then fail.
- Choose Save and apply.
Important
The form has no field for a CA certificate, and saving from it clears one set for the target before. The connection test does not use a CA certificate either, so a store whose certificate is signed by a private CA fails it unless the Lifecycle module already trusts that CA.
Saving tests the store once more, saves the target in the platform configuration, and starts a rollout of the Velero dependency that renders it from the new target; the page takes you to the Rollouts tab of the Lifecycle Manager. The change is refused, and the previous target stays, when:
- the store refuses the credentials or cannot be reached;
- another rollout is executing;
- someone else changed the platform configuration in the meantime — reload and try again;
- a pre-upgrade check for that rollout fails — fix it first, as it cannot be overridden from here;
- the platform has no saved configuration to change, on a platform installed before the installer saved it. Edit target is then unavailable; re-run the installer to save it.
Warning
Pointing at another bucket or endpoint leaves the earlier backups behind, and it moves every database's archive — tenants' databases included — to the new store. The earlier backups are no longer listed here, and a database can only be restored to a point before the switch from the old store, so keep the old bucket for as long as those backups are wanted. The rollout takes a fresh base backup of the platform database in the new store once its archive has moved there. A tenant's database gets none: it cannot be restored from the new store until its next base backup there has completed.
Change the daily schedule
The Daily schedule card shows when the backup runs, the next run, the last backup and how long each backup is kept, with Active, Paused or Missing.
- Edit schedule changes the Cron expression (UTC) — five fields — and Keep each backup for (days), at least one day. Velero deletes a backup and its data once it is older than that. The retention applies to backups taken after the change, Back up now included; one already taken keeps the date in its Expires column.
- Pause the daily backup turns the schedule off and on. No daily backup runs while it is paused, and Backup health says so.
Important
Every rollout of the Velero dependency — including the one that changing the target starts — writes the schedule again with the installer's values: 02:00 UTC and 30 days. A changed cron expression or retention is lost then and has to be set again; a pause stays.
A Missing schedule is written again by a rollout of the Velero dependency; until then the card cannot be edited and Back up now is refused.
A database is filling up behind a failing archive
When a database cannot archive its write-ahead log, the page shows A database is filling up behind a failing archive, naming the database and, where it can, how long it has before it stops accepting writes.
- Fix the target first. The archive writes to the backup target, and a database catches up on its own once the target is reachable again.
- If the deadline is closer than the fix, choose Suspend archiving for the database and confirm. The database can then recycle its write-ahead log and keeps accepting writes, but everything written from then until archiving resumes cannot be recovered.
- When the target works again, choose Resume archiving for the database. Archiving starts again and a new base backup is taken; the warning stays until that backup has completed.
Suspending is only offered while a database's archiving is failing, and resuming only for a database suspended from this page or through the API.
Caution
The page covers every Postgres database on the cluster that archives to the backup target, tenants' databases included. Suspending archiving for a tenant's database costs that tenant the same unrecoverable window.
Through the API
For example with stsh api:
stsh api POST /lifecycle/backups
stsh api GET /lifecycle/backups/conditions
stsh api PUT /lifecycle/backups/schedule -d '{"cron": "0 2 * * *", "ttlHours": 720}'
stsh api PUT /lifecycle/backups/schedule -d '{"paused": true}'
stsh api POST /lifecycle/backups/archiving/suspend -d '{"namespace": "<namespace>", "name": "<cluster>"}'Every route is listed in the module's API reference.