Platform backups
What the platform backs up every day, where the backups go, which backups the Backups page lists, and what its health conditions mean.
The platform backs itself up with Velero, and its own database additionally with CloudNativePG.
Platform administrators see and manage both under Settings → Backups, which also appears on
the Admin page. Opening it needs lifecycle/view at the platform root.
What is backed up
- The daily platform backup. The Velero schedule
platform-dailybacks up every namespace exceptkube-system,kube-public,kube-node-leaseandvelero, with the cluster-wide resources and the volumes — tenants' resources included, so a lost cluster can be rebuilt with them. The installer sets it to run at 02:00 UTC and keeps each backup for 30 days. - The platform database. The platform's Postgres cluster archives its write-ahead log to the backup target continuously and takes a daily base backup. A tenant's Postgres database with backups turned on archives to the same target, under a path of its own.
- Before a rollout that upgrades a third-party chart, the operator takes a snapshot of its own — see Upgrade the platform.
- Back up now takes an extra backup with the same scope as the daily one.
The page takes and lists backups; it does not restore them.
Where the backups go
Every backup goes to the target: the bucket of Velero's default storage location, on AWS S3 or an S3-compatible store. The Target card shows its provider, bucket, prefix, region, endpoint, access key id, when it was last validated, and whether it is available.
Keep the target outside the cluster it backs up. A target whose endpoint is a service inside this cluster is flagged: a failure that takes the cluster takes the backups with it.
The backup list
The table lists the platform's own backups, newest first: the daily ones, the pre-rollout snapshots, the ones taken with Back up now, and the platform database's backups. A tenant's snapshots taken by hand and its databases' backups are not in it, but the backups that a backup policy on a tenant's resource takes are, with Rollout as their source. Each row shows the Phase Velero or CloudNativePG reports, Started, Duration, Expires, the Source — Schedule, Rollout or Manual — and Problems, with the details behind Show details.
Backup health
Backup health judges four conditions, each OK, Warning, Failing or Unknown:
| Condition | Warning or Failing when |
|---|---|
| The backup target is available | Failing: there is no default storage location, or it is not Available — every backup fails until it is. Warning: its endpoint is inside this cluster |
| The platform is backed up daily | Failing: the platform-daily schedule is missing, or its newest backup failed. Warning: the schedule is paused, has not produced a backup yet, or its newest backup started more than two days ago |
| Every backup policy is producing backups | Warning: a resource backup policy's last run failed, or it has not succeeded within twice its interval |
| No database is filling up behind a failing archive | When a Postgres database — the platform's or a tenant's — cannot archive its write-ahead log, it cannot recycle it either, so its volume fills until it stops accepting writes. The condition names the database and, from the measured fill rate, how long it has left. Failing when that is under two hours or less than a tenth of the volume is free; Warning otherwise, and while archiving is suspended or a resume's base backup has not completed |
Unknown means the Lifecycle module was not allowed to read the schedule or the backup policies. The other two conditions have no Unknown. The platform also raises a platform alert, one per database, when a database fills up behind a failing archive or has its archiving suspended. It checks every five minutes, so the alert reaches someone even while nobody has the page open.
To change the target or the schedule, or to act on a filling database, see Manage platform backups.