Upgrade the platform
Check for new releases, plan a rollout, review its batches and pre-upgrade checks, execute it, follow it, and deal with a pause, a cancellation or a rollback.
Requires: lifecycle/plan, lifecycle/execute
Planning a rollout needs lifecycle/plan and executing it lifecycle/execute, both at the
platform root; rolling one back needs lifecycle/rollback. Platform Owner and Platform
Contributor hold all three. Everything below happens under Settings → Lifecycle Manager.
Check for new releases
The platform reads the release feed about once an hour. To read it now, open the Components tab and choose Check for updates. The answer is one of:
- Found N new release(s) — the Version column now shows the newer release next to the running version, and the line above the table counts the components that can be upgraded.
- No new releases — everything is up to date — the usual answer.
- Could not reach the release feed — the feed could not be fetched, or its signature could not be verified. Nothing changed. Check that the platform can reach the feed; if it can, see How the platform reads the feed.
- No release feed configured — this platform reads no feed, so there is nothing to check.
Checking only records which releases exist; it never changes what runs.
Plan a rollout
Choose New rollout, in the page header or on the Rollouts tab. The planning page lists the components that have a newer release (Outdated); All lists every component, and Filter components narrows the list by name.
- Choose Select outdated to tick every outdated component with its newest release as the target — the CRD bundle included, because the plan can only put the CRDs first if they are in it. Or tick components one by one. Clear unticks everything.
- For each ticked row, pick the Target version. The list shows each release's channel and, for a CRD bundle, the installer that release runs. To roll a component back by hand, find it under All and pick an older release.
- Choose Review plan.
The columns are Component, Kind, Installed (what runs now), Rolls after (the components it depends on) and Target. A few things are decided for you:
- Components you did not tick stay as they are. A component can roll on its own: what it depends on only decides the order when both are in the plan.
- A third-party chart that a ticked module pins appears indented under that module and rolls with it. Untick it to leave it behind; the plan then says so.
- A chart marked Manual upgrade is upgraded by hand through its vendor's procedure and cannot be ticked.
- The Platform step (installer) is not in the list. Any plan that moves the CRD bundle rolls a platform release, and the platform adds that release's installer as a step of its own.
Review the plan
The review replaces the list, and the page's address now carries the plan, so a reload or a colleague with the link sees the same plan. Edit selection goes back to the list and drops the plan.
Read it from the top:
- Batches — each batch finishes before the next starts, and the components in one batch roll together. Third-party charts come first, one per batch; then the CRD bundle; then the platform step; then the Lifecycle module on its own when it rolls with others, because its new version renders the configuration of the rest; then everything else in dependency order. Each row shows the version it rolls from and to.
- Dependencies — for each third-party chart, the result of a dry run of its upgrade once the plan is validated, whether undoing it would mean restoring a database, and whether the vendor only supports one minor version at a time.
- Left out of this plan — what the plan could not include and why, for example a chart a module pins that cannot roll yet, or a release whose platform step cannot run. The rest of the plan goes ahead without it.
The side panel shows the plan's status. Validated means it can be executed. Any problem listed in red above the batches keeps the plan unvalidated, and Execute rollout does not appear. Typical problems: a target version no release carries; a release that is marked schema-breaking; a release that may only be applied from a newer installed version; a chart that would skip a minor version its vendor requires; or a release that also ships a CRD bundle that has not been applied, while the bundle is not in the plan. Change the selection and review again.
Deal with the pre-upgrade checks
A validated plan runs the pre-upgrade checks, and the side panel counts them as Passed, Warnings, Failed and Skipped. The full report is below the plan; every check is described in Pre-upgrade checks.
- A warning never stops the rollout.
- A failed check has to be overridden deliberately: tick Roll out anyway, overriding the failed checks. The override, and the ids of the checks it covered, are recorded on the rollout.
- A failed check marked Cannot be overridden keeps Execute rollout unavailable, and the tick box leaves it out: the rollout would only be refused later, at the step that needs what is missing. The most common is the platform admin bundle, which a cluster administrator applies with the command the check shows — see The platform admin bundle. Once it has been fixed, choose Run checks again.
Re-run checks runs the report again after you have fixed something.
Execute the rollout
Choose Execute rollout. Before anything changes, the platform:
- refuses when another rollout has not ended — only one runs at a time, and a paused rollout counts until it ends;
- refuses when the plan upgrades a third-party chart or runs a platform step and the running operator is too old to carry that out;
- runs the pre-upgrade checks again and decides on that run, not on the one you reviewed, so a check that has failed since stops the rollout unless you overrode it;
- writes the configuration the release gives each component, and refuses with the reason when a value that configuration needs is missing on this platform.
When it starts, the page takes you to the rollout's own page.
Follow the rollout
The rollout's page has an address of its own; send it to anyone who asks how the upgrade is going. Every rollout is also listed on the Rollouts tab, with Details to open it.
- Progress shows each batch and, per component, whether it is queued, rolling or done, the version, and the image digest it rolls from and to. A third-party chart shows its Helm revisions instead.
- Status, beside it, shows the rollout's State and Phase, when it started and who started it, and holds the actions.
- Pre-rollout snapshot appears for a rollout that upgrades a third-party chart: the backup taken before batch 1.
- Checks after dependency batches appears when checks are repeated after a chart batch.
- The pre-upgrade checks are shown as they stood when the rollout started, with any override.
What the states mean is listed in Statuses.
When a rollout pauses
A rollout pauses rather than carrying on when a component fails to roll out — its Deployment is missing or stops progressing, or its chart upgrade, CRD bundle or platform step is refused —, when a component fails its health check, when a check repeated after a chart batch turns worse, when the cluster refuses the operator something the rollout needs, when its pre-rollout snapshot fails, or when its platform step finds platform settings that were changed by hand. The page says so at the top, with the reason, and the failing component shows the detail.
- Resume — after you have fixed the cause. The paused batch is checked again and the rollout continues.
- Resume without snapshot — offered only when the snapshot failed and may be skipped. It continues without a backup of the upgraded charts' resources, and the override is recorded. A chart whose undo is a database restore never offers it.
- A pause on settings changed by hand offers Revert changes as well — see Settings changed by hand.
A batch that does not finish within its time limit — ten minutes, or thirty for a batch that upgrades a third-party chart — ends the rollout as Failed instead of pausing it. A resumed batch starts a new time limit.
Cancel a rollout
Cancel rollout ends a rollout that has not finished. It ends as Cancelled; components that were already upgraded stay on their new version. Cancelling stops the rollout from taking its next step, but does not stop a change already under way: a component of the current batch whose new version was already applied goes on rolling out, and a chart upgrade or platform step that has started runs to its end.
Roll back
Roll back is offered on a rollout that has ended — succeeded, failed or cancelled — and
needs lifecycle/rollback. It starts a new rollout that takes every upgraded component back to
the image it ran before, in reverse order, and takes you to that rollout. Failed pre-upgrade
checks do not stop it: a rollback is how you recover from an unhealthy platform.
- The CRD bundle and the platform step never roll back. CRD changes only ever add, and older components ignore what was added.
- A third-party chart returns to the Helm revision it had before the upgrade. A chart that migrates its database forward is left out: undoing it means restoring the platform database from the pre-rollout backup and then rolling the chart back by hand.
- A module that migrated its own database forward may not run on its previous image. The confirmation says so.
After the rollout
Open the After upgrading tab. A release can leave a step only a person can take, and a count on the tab and a reminder at the top of every page say when it has — see After upgrading. After a rollout succeeds, the platform also scans for settings changed by hand and, for a release rollout, records the release as the platform's version. It does both once it notices the success: at once while the rollout's page is open, otherwise the next time anything checks whether a rollout is running — at the latest with the next scheduled scan, within six hours.
Through the API
The same flow is available through the API, for example with stsh api:
stsh api GET /lifecycle/components
stsh api POST /lifecycle/rollouts/plans -d '{"targetVersions": {"apps": "<version>"}}'
stsh api POST /lifecycle/rollouts/plans/<plan-id>/validate
stsh api GET "/lifecycle/preflight?planId=<plan-id>"
stsh api POST /lifecycle/rollouts/executions -d '{"planId": "<plan-id>", "acknowledgedFailures": []}'
stsh api GET /lifecycle/rollouts/executions/<execution-id>List a failed check's id in acknowledgedFailures to override it. Every route is listed in the
module's API reference.