Error codes
The codes the Lifecycle module refuses a request with — platform drift and platform backups — what each means, and what to do about it.
Every route below is under /lifecycle. The portal turns these codes into its own messages; this
page is for reading them in an API response, for example from stsh api. A refusal before a route
runs keeps the platform-wide codes — a missing permission is FORBIDDEN — see
Status codes.
Two body shapes
The module answers in two shapes, and the code is in code in both.
Platform drift codes are in PascalCase, in the platform's usual error body — see The error body. Extra fields sit at the top level:
{
"error": "A platform drift scan is already running.",
"code": "ScanRunning",
"module": "lifecycle",
"correlationId": "3f2b9c1e-6d4a-4e0b-9a51-2c7d8e0f4b6a",
"scanId": "8d0c6f4e-2a71-4b4c-9d55-6f1e0b7a2c93"
}Platform backup codes are in kebab-case, in a problem details body
(application/problem+json). The code is also the title, the sentence for a person is in
detail, and there is no error field:
{
"type": "https://tools.ietf.org/html/rfc9110#section-15.5.10",
"title": "backup-in-progress",
"status": 409,
"detail": "backup platform-manual-20260927-081500 is InProgress; Velero runs backups one at a time, so a second request would only queue behind it",
"code": "backup-in-progress",
"backup": "platform-manual-20260927-081500",
"phase": "InProgress",
"module": "lifecycle",
"correlationId": "3f2b9c1e-6d4a-4e0b-9a51-2c7d8e0f4b6a"
}Platform drift
What these refusals protect is described in Settings changed by hand. They come from three routes:
| Route | What it does |
|---|---|
POST /lifecycle/platform-drift/scan |
Starts a scan — Check now |
POST /lifecycle/platform-drift/adopt |
Adopts a change into the configuration, with scanId, configKey, liveValue and acknowledge |
POST /lifecycle/rollouts/executions/{id}/revert-platform-drift |
Puts the release's values back on a paused rollout, with acknowledge and optionally fields |
AcknowledgementRequired
400 on adopt and revert. The body did not carry "acknowledge": true; a request with no body
counts as not acknowledged. On revert this is checked before the rollout is looked up, so an
unknown rollout id without the acknowledgement gets this code rather than 404.
Send the request again with "acknowledge": true, having read what it changes: adopting rewrites
the platform configuration, and reverting loses what was changed by hand.
NotPausedOnPlatformDrift
409 on revert. The rollout is not paused on platform drift: it is running, finished, or paused for another reason. The message names its phase and the reason for the pause.
There is nothing to revert. Resume or cancel the rollout as its page offers, or reload it to see its current state.
OperatorCapabilityMissing
409 on revert. The operator running in the cluster is too old to overwrite changes made by hand.
Keep the change or undo it yourself, then resume the rollout — see When a release rollout pauses on it. Upgrading the operator makes revert available for later rollouts.
PlatformDriftChanged
409 on revert. The changes in the cluster are no longer the ones you reviewed: the fields you
sent do not match the current changes, a changed component is no longer in the batch the rollout is
paused on, or the cluster refused the revert because the object moved while it was being made.
The body carries platformDrift, the current changes, each with component, step, kind,
namespace, name, field, manager and message. Review them and send the revert again with
them as fields.
ScanRunning
409 on scan. A scan is already running. The body carries its scanId.
Wait for it to finish, then read the result with GET /lifecycle/platform-drift.
RolloutBusy
409 on scan and adopt. A rollout is executing or paused; a paused rollout counts until it is resumed to the end, cancelled or rolled back.
Finish or cancel the rollout first. After an adopt refused this way, wait for the scan that follows the rollout and review its result before adopting.
ScanStale
409 on adopt. The scanId you sent is missing or is not the latest finished scan, a scan is
running now, or no scan has finished yet.
Read the latest scan with GET /lifecycle/platform-drift, review it, and adopt with its scanId.
DriftChanged
409 on adopt. The liveValue you sent is missing or is not the value the latest scan found in
the cluster.
Review the latest scan again and adopt the value it shows.
ConfigChanged
409 on adopt. The platform configuration, the Secret stackship-install-config, has changed
since the scan read it, no longer holds the value the scan saw, or is missing. It is also returned
when another change is written to the configuration at the same moment. Any change to the
configuration counts — changing the platform backup target is one.
Start a scan with Check now, review its result, and adopt again.
NotAdoptable
404 on adopt. The latest scan offers no adoption of the configKey you sent: the change cannot
be adopted, or the key is not a platform setting (platform. followed by its name) whose value is
text. This is checked before liveValue and before any running rollout.
The scan's result says why a change cannot be adopted — see
When a change cannot be adopted. Let a release rollout put the
value back, undo the change by hand, or write it into stackship-install-config yourself.
Platform backups
What these routes do is described in Manage platform backups.
schedule-missing
409 on POST /lifecycle/backups and PUT /lifecycle/backups/schedule. The Velero schedule
platform-daily does not exist. A manual backup takes its scope from it, so none can be taken, and
there is no schedule to change. The schedule is checked before the request is validated.
Roll out the Velero dependency, which writes the schedule again — see Change the daily schedule.
backup-in-progress
409 on POST /lifecycle/backups. A backup taken with Back up now, or through this route, is
still queued or running and was started less than three hours ago. A running daily backup does not
cause it. The body carries backup, the name of that backup, and phase, its Velero phase or
queued.
Wait for that backup to finish. One that is stuck stops counting three hours after it started — see Take a backup now.
invalid-target
400 on PUT /lifecycle/backups/target. The target is not valid: the bucket, access key id or
secret access key is missing, the endpoint is not an absolute http or https URL, a provider
other than aws is named, or the CA certificate holds anything but valid PEM certificates.
detail says which.
Correct the request and send it again.
target-rejected
422 on PUT /lifecycle/backups/target. Listing the bucket with the new credentials failed: the
store refused them, or could not be reached within ten seconds. detail carries the store's answer.
Nothing was saved.
Correct the credentials, endpoint or bucket. POST /lifecycle/backups/target/test runs the same
test without saving and never answers with an error code — see
Change the backup target.
target-not-applied
409 on PUT /lifecycle/backups/target. The new target could not be applied, for one of several
reasons that only detail tells apart:
- The platform has no saved configuration: no
stackship-install-config, or one without a configuration in it, on a platform installed before the installer saved it. Re-run the installer to save it. - The saved configuration cannot be read as an install request. Re-run the installer.
- The Velero dependency cannot be rolled out again from the new target, or the rollout that applies
it did not validate;
detailcarries the reason or the plan's warnings. - The previous target could not be written back after the change was refused. The saved target then
disagrees with the cluster: correct
stackship-install-configby hand.
Except in the last case, the previous target stays.
rollout-in-flight
409 on PUT /lifecycle/backups/target. Another rollout is executing or paused, so the rollout
that applies the target cannot start. The previous target stays.
Let the rollout finish, or cancel it, and save the target again.
target-changed
409 on PUT /lifecycle/backups/target. The platform configuration was written by someone else
between reading it and saving the new target. Any change counts, adopting a changed setting
included. Nothing was saved.
Reload the target and save it again.
preflight-failed
409 on PUT /lifecycle/backups/target. A pre-upgrade check of the rollout that applies the
target failed. The failed checks are named in detail only. A failure cannot be overridden from
this route, and the previous target stays.
Fix what the check reports — see Pre-upgrade checks — and save the target again.
invalid-schedule
400 on PUT /lifecycle/backups/schedule. cron is not a five-field cron expression, or
ttlHours is less than 1.
Correct the request and send it again.
invalid-cluster
400 on POST /lifecycle/backups/archiving/suspend and …/resume. The body does not name the
database cluster: namespace or name is empty.
Send both.
cluster-not-found
404 on POST /lifecycle/backups/archiving/suspend and …/resume. There is no database cluster
with that namespace and name.
Check both; the warning on the Backups page names the database.
archiving-not-configured
409 on POST /lifecycle/backups/archiving/suspend. The database does not archive to the backup
target, so there is nothing to suspend.
archiving-healthy
409 on POST /lifecycle/backups/archiving/suspend. The database is not filling up behind a
failing archive, or its archiving is already suspended. The message says the database is not under
archive pressure in both cases.
Suspending opens a window that cannot be recovered, so it is refused unless the archive is failing. If the archiving is already suspended, resume it once the target works — see A database is filling up behind a failing archive.
archiving-not-suspended
409 on POST /lifecycle/backups/archiving/resume. The database's archiving was not suspended
from the Backups page or this API, and no earlier resume is still waiting for its base backup.
There is nothing to resume.