A rolling update can temporarily run old and new replicas together. The scheduler admits the new Pods from requests, while the node must survive their actual peaks plus system reservations. If a large service is already close to node memory capacity, maxSurge can leave new Pods Pending or push the node into pressure. Setting maxUnavailable higher may free capacity, but it can cut serving capacity before the new version proves ready.
Memory-safe rollouts: reserve space for old and new Pods
Operational decision
A reporting API runs six replicas, each requesting 896 MiB, on nodes whose remaining allocatable memory is only 1.3 GiB. A rollout with two surge Pods needs 1792 MiB of additional requests before either old Pod exits, so it cannot fit on that pool even if current usage is low. Add node capacity or use a smaller surge after verifying the service can meet its availability target; do not solve the scheduling block by falsifying requests. Test cold-start memory and readiness under load, then drain an old Pod and inspect node pressure, restart counts, and p99 latency. Include sidecars and DaemonSet overhead in the headroom ledger.
Reporting API rollout
Replicas: 6
Request per new Pod: 896 MiB
maxSurge: 2
Additional requested memory: 1792 MiB
Available node allocatable headroom: 1331 MiB
Decision: add capacity or revise rollout after availability testCost and verification
The scheduling comparison is O(P) over pending Pods for a simple capacity ledger, but real placement also depends on affinity, topology, quotas, and per-node fit. One aggregate free-memory number is insufficient if no single node can host a Pod. Extra surge capacity costs node time, while low surge increases rollout duration. Measure schedule-to-ready time, unavailable endpoints, MemoryPressure events, and user latency for the chosen strategy.
Common Mistakes
- Do not use average free memory as proof a Pod fits on a node.
- Do not lower memory requests merely to force scheduling.
- Do not increase maxUnavailable without checking serving capacity.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Kubernetes Deployment: rolling update capacity
- Topology spread: keep replicas out of one failure domain
- Node autoscaling: make pending Pods schedulable before traffic rises
