Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Linux kernel rollouts: prove the running kernel after each reboot cohort

Last updated: 5 Oct 20267 min read
tutorial
AdvancedBy AITrove Editorial

A kernel rollout has three states: the package available on disk, the boot entry selected for the next start, and the kernel running now. Only the last state proves the host crossed the reboot boundary. A live patch can alter selected running functions without a reboot, but its coverage is narrower than replacing the running kernel and it does not remove the need to plan later boot transitions. Treat a fleet patch as a sequence of host cohorts with a maximum unavailable count, health gates, and a known prior boot entry. Preserve the exact package and bootloader state before changing either.

Operational decision

A 93-node API fleet needs a kernel security update. The operator drains one canary node, records its current release and boot entry, installs the approved package, and reboots it through the normal maintenance path. After boot, they compare uname output with the approved release, check failed units, mount and network state, and run a synthetic API write through the load balancer before admitting traffic. A second cohort of five starts only after the canary remains healthy for an observation period. If the new kernel fails before remote access returns, a console or provider rescue path selects the retained prior entry. If the host boots but the API path fails, rollback includes draining it again and proving the earlier kernel actually runs after reboot.

bash
uname -r
cat /proc/cmdline
systemctl --failed

Cost and verification

Serial cohorts extend the maintenance window, but a fleet-wide reboot risks losing all replicas to the same driver or boot error. Measure time from drain to application acceptance, count hosts running the intended release, and alert on a host stuck between installed and running versions. Rollback is not a package-manager command alone; the selected boot entry and service health must be checked again. Keep workload quorum and capacity margins during each cohort. If the service has local state, confirm its shutdown and replay contract before draining.

Common Mistakes

  • Do not mark a host patched because a package manager returned success.
  • Do not assume a live patch covers every kernel change or survives an unrelated reboot plan.
  • Do not delete the earlier boot entry before the canary and recovery path are proven.

Connected lessons

Practice and check

devops
linux
host-security
Storage details