A node with low CPU use is not automatically removable. Its Pods may have large resource requests, a restricted zone, local storage, scarce attach slots, or a PodDisruptionBudget that currently allows no voluntary eviction. A consolidation controller that respects the Eviction API can wait for those constraints, while an involuntary node loss remains possible. Cost savings require a post-move node count, not just an eviction attempt.
Node consolidation: calculate the capacity needed to evict safely
Operational decision
A parcel API has one lightly used node beside a stateful ledger and a scheduled batch worker. Before consolidating, list requests, node affinity, topology spread, volume attachments, PDB disruptions allowed, and the surge reserve for the next release. The read-only commands below are a starting inventory; test eviction and replacement in an isolated pool. Move a replayable worker first, verify the queue lease and workload latency, and only then attempt the API Pod. If no destination node has enough allocatable resources, consolidation will fail or trigger a new node, producing no saving. Check how the provider bills partial node-hours and any minimum node-pool size. Keep capacity for a node failure and rollback after consolidation. If a PDB blocks a voluntary eviction, adjust workload availability or schedule maintenance; do not bypass it with direct Pod deletion to obtain a short-lived saving. Record exactly which node-hours were removed and which replacement resources were added.
kubectl -n parcel get pdb -o wide
kubectl -n parcel get pods -o wide
kubectl get nodes -L topology.kubernetes.io/zone
kubectl -n parcel describe pod parcel-api-canaryCost and verification
Consolidation may reduce node-hours but increase scheduling latency, hotspot risk, or recovery time. A drain also creates temporary duplicate capacity while replacements start, so the billing benefit may begin later than the controller's decision. Measure actual nodes removed, replacement starts, disruptions allowed, pending-Pod time, p95 latency, and recovery headroom. Count repeated failed attempts as control-plane and operator cost even when they do not change the bill.
Common Mistakes
- Do not classify a node as removable from CPU utilization alone.
- Do not bypass PDBs or local-volume constraints to force a cheaper topology.
- Do not report savings until the billed node count actually falls.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Pod disruption budgets: make node drains measurable
- Node autoscaling: make pending Pods schedulable before traffic rises
- CSI attach limits: verify storage placement as well as CPU placement
- Shared cluster cost: reconcile service allocation with the provider bill
