Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Fleet promotion: bound the number of clusters changed at once

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

Fleet promotion groups targets into explicit cohorts with maximum concurrency, stable selection rules, health and user-path gates, and a stop action. A controller's creation order or a branch update is not a safety policy. Review the exact target set for each wave; label changes, missing selectors, or an unselected cluster can alter coverage. Distinguish application health from business acceptance, and define how to hold later cohorts without rolling back already healthy ones.

Operational decision

A fulfillment platform ships a new sidecar to 27 clusters. The first cohort contains one low-traffic cluster, the second has three diverse environments, and the remainder is split by region and capacity. Before release, generate the target matrix from current cluster inventory and compare it with the intended set. A promotion record pins the same rendered input and image digest for every cohort; any cluster-specific overlay difference is listed separately. Advance only after the previous group reports observed revision, healthy workloads, and successful synthetic fulfillment. If one cluster fails, stop new syncs, preserve its logs and render, and inspect whether the failure follows architecture, policy, or local dependency. A later cohort must not start simply because the first group's controller status changed to Healthy for a moment. Test the stop path in an isolated fleet and include a cluster with an intentionally stale source artifact. Record clusters that were not selected so the rollout does not end with a silent old-version island. When recovery begins, use the same cohort limits to avoid multiplying rollback load across every cluster at once.

Output
Fulfillment fleet contract
Target inventory: 27 named clusters at release start
Cohorts: 1, then 3, then bounded regional groups
Per-cluster: source, render, image, observed revision
Advance: workload health plus synthetic fulfillment
Stop: failed group blocks unsynced targets
Exception: unselected or stale cluster requires disposition
Rollback: bounded groups with user-path proof

Cost and verification

For K clusters and G cohorts, status collection is O(K) per polling pass; too-frequent probes multiply API and application load. Smaller cohorts reduce blast radius but increase total release time, especially when each gate waits for a long observation window. Measure maximum concurrent changes, time to first failure, unselected targets, source age, and customer errors by cohort. A fleet-wide rollback should be capacity tested, because it can stress registries and schedulers more than the original rollout.

Common Mistakes

  • Do not use alphabetical cluster order as a substitute for risk-based cohorts.
  • Do not let a missing selector silently exclude a target.
  • Do not advance a group on a momentary health status without user-path evidence.

Connected lessons

Practice and check

Platform operating-contract follow-up

devops
gitops
Storage details