Skip to content
Stackship documentation Svenska

Lifecycle ManagerAdministrators

Manage platform backups

Take a platform backup now, change where backups go, change or pause the daily schedule, and keep a database from filling up behind a failing archive.

Requires: lifecycle/execute, lifecycle/backupsManage

Everything here is under Settings → Backups. Back up now needs lifecycle/execute at the platform root, which Platform Owner and Platform Contributor hold. Changing the target or the schedule, and suspending or resuming archiving, needs lifecycle/backupsManage, which of the platform roles only Platform Owner holds: a wrong target loses the way back for the whole platform.

Take a backup now

Choose Back up now. The platform starts a backup named platform-manual-<date>-<time>, with the same scope as the daily backup, and it appears in the list.

Velero runs backups one at a time, so a new one is refused while a backup taken with Back up now is still queued or running; one started more than three hours ago no longer counts. A daily backup that is running does not refuse it — the new one queues behind it. It is refused as well when the platform-daily schedule is missing, because a manual backup takes its scope from it.

Change the backup target

  1. On the Target card, choose Edit target.
  2. Fill in Bucket, Region, Endpoint, Checksum algorithm, Access key id and Secret access key. The form starts from the current target, except the secret key, which is never shown again after it is saved.
    • Endpoint — leave it empty for AWS S3; for MinIO, Ceph or another S3-compatible store, the full https:// URL.
    • Region — S3-compatible stores usually accept any region.
    • Checksum algorithm — leave it empty for MinIO releases older than mid-2024, SeaweedFS and Ceph RGW, which reject the default checksums.
  3. Choose Test connection. The store is asked to list the bucket with these credentials, and its answer is shown. Listing is all it checks: a key that may read the bucket but not write to it passes, and the backups then fail.
  4. Choose Save and apply.

Important

The form has no field for a CA certificate, and saving from it clears one set for the target before. The connection test does not use a CA certificate either, so a store whose certificate is signed by a private CA fails it unless the Lifecycle module already trusts that CA.

Saving tests the store once more, saves the target in the platform configuration, and starts a rollout of the Velero dependency that renders it from the new target; the page takes you to the Rollouts tab of the Lifecycle Manager. The change is refused, and the previous target stays, when:

  • the store refuses the credentials or cannot be reached;
  • another rollout is executing;
  • someone else changed the platform configuration in the meantime — reload and try again;
  • a pre-upgrade check for that rollout fails — fix it first, as it cannot be overridden from here;
  • the platform has no saved configuration to change, on a platform installed before the installer saved it. Edit target is then unavailable; re-run the installer to save it.

Warning

Pointing at another bucket or endpoint leaves the earlier backups behind, and it moves every database's archive — tenants' databases included — to the new store. The earlier backups are no longer listed here, and a database can only be restored to a point before the switch from the old store, so keep the old bucket for as long as those backups are wanted. The rollout takes a fresh base backup of the platform database in the new store once its archive has moved there. A tenant's database gets none: it cannot be restored from the new store until its next base backup there has completed.

Change the daily schedule

The Daily schedule card shows when the backup runs, the next run, the last backup and how long each backup is kept, with Active, Paused or Missing.

  • Edit schedule changes the Cron expression (UTC) — five fields — and Keep each backup for (days), at least one day. Velero deletes a backup and its data once it is older than that. The retention applies to backups taken after the change, Back up now included; one already taken keeps the date in its Expires column.
  • Pause the daily backup turns the schedule off and on. No daily backup runs while it is paused, and Backup health says so.

Important

Every rollout of the Velero dependency — including the one that changing the target starts — writes the schedule again with the installer's values: 02:00 UTC and 30 days. A changed cron expression or retention is lost then and has to be set again; a pause stays.

A Missing schedule is written again by a rollout of the Velero dependency; until then the card cannot be edited and Back up now is refused.

A database is filling up behind a failing archive

When a database cannot archive its write-ahead log, the page shows A database is filling up behind a failing archive, naming the database and, where it can, how long it has before it stops accepting writes.

  1. Fix the target first. The archive writes to the backup target, and a database catches up on its own once the target is reachable again.
  2. If the deadline is closer than the fix, choose Suspend archiving for the database and confirm. The database can then recycle its write-ahead log and keeps accepting writes, but everything written from then until archiving resumes cannot be recovered.
  3. When the target works again, choose Resume archiving for the database. Archiving starts again and a new base backup is taken; the warning stays until that backup has completed.

Suspending is only offered while a database's archiving is failing, and resuming only for a database suspended from this page or through the API.

Caution

The page covers every Postgres database on the cluster that archives to the backup target, tenants' databases included. Suspending archiving for a tenant's database costs that tenant the same unrecoverable window.

Through the API

For example with stsh api:

bash
stsh api POST /lifecycle/backups
stsh api GET /lifecycle/backups/conditions
stsh api PUT /lifecycle/backups/schedule -d '{"cron": "0 2 * * *", "ttlHours": 720}'
stsh api PUT /lifecycle/backups/schedule -d '{"paused": true}'
stsh api POST /lifecycle/backups/archiving/suspend -d '{"namespace": "<namespace>", "name": "<cluster>"}'

Every route is listed in the module's API reference.