Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Cluster upgrade drill: preserve a path through each version step

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

A Kubernetes cluster upgrade changes API server, controllers, kubelets, add-ons, and sometimes the container runtime at different times. The supported version skew between components constrains the order of operations; the exact rule depends on the versions and distribution in use. A successful control-plane upgrade does not prove that applications survive node replacement or that old manifests still apply.

Operational decision

For a two-zone claims cluster, inventory deprecated APIs and controller versions, then rehearse the target version in a test cluster with representative workloads. Check PDB disruptionsAllowed and spare placement capacity before draining each node. Upgrade the control plane using the provider-supported path, then move one node pool at a time while watching request errors, pending Pods, storage attach, and ingress routes. The runbook outline is deliberately provider-neutral; use the distribution's procedure for the actual command sequence. If an add-on fails, stop before advancing another pool. Recreate a node from a known image when required; do not assume a control-plane downgrade is supported. Keep a service-level recovery plan for workloads that cannot roll back their infrastructure in place.

Output
Claims cluster upgrade gate
1. Record current and target versions and supported skew.
2. Inventory removed APIs and add-on compatibility.
3. Rehearse control plane plus one node pool in test.
4. Check PDBs, zone headroom, backups, and user-path probes.
5. Upgrade one production pool; measure errors and pending Pods.
6. Stop or proceed on the recorded service thresholds.

Cost and verification

Extra node-pool capacity and a test cluster cost money during the upgrade window. Without them, a drain can stall or evict more useful capacity than the application tolerates. Cluster control-plane health, workload availability, and data integrity need separate gates. Version-skew guidance can change between releases, so review the target release's current support policy and provider limits when scheduling the work.

Common Mistakes

  • Do not drain nodes before checking PDBs and placement headroom.
  • Do not treat a green control plane as proof that user requests work.
  • Do not assume the cluster can be downgraded after a failed upgrade.

Connected lessons

Advanced follow-up

Advanced follow-up

devops
resilience
Storage details