A topology spread constraint tells the scheduler how evenly matching Pods should occupy domains such as zones or nodes. The topology key names a node label, the selector identifies the peer Pods, and maxSkew limits the accepted imbalance. With DoNotSchedule, an unsatisfied rule leaves a new Pod pending; ScheduleAnyway expresses a preference and may permit an uneven placement.
Topology spread: keep replicas out of one failure domain
Operational decision
A four-replica claims API serves two zones. Start with one zone constraint and inspect node labels before rollout. The sample belongs in a Deployment Pod template, not at the top level of a Deployment. It assumes both zones have enough allocatable CPU and memory. Test a node loss and a zone loss separately: spreading four replicas across zones does not ensure two nodes per zone. Pair placement with resource requests, replica count, a PDB, and traffic routing. If a new revision stays pending, inspect scheduler events and domain eligibility before weakening the rule. Cross-zone calls and storage placement may change latency and network cost.
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels: {app: claims-api}Cost and verification
Hard placement constraints consume headroom because free capacity in one zone cannot always compensate for a full second zone. A label missing from a node changes which domains are eligible. Measure pending time after autoscaler events and during rollout, not only after steady state. A spread rule improves failure isolation but adds scheduling work and may increase cross-zone traffic; the service's recovery goal decides whether that trade is justified.
Common Mistakes
- Do not use a selector that misses the actual workload label.
- Do not call zone spread a node-level guarantee.
- Do not add DoNotSchedule without checking capacity in every eligible domain.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Pod disruption budgets: make node drains measurable
- Kubernetes requests and limits: schedule for real load
- Multi-region failover: define write ownership before moving traffic
