Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Node consolidation: calculate the capacity needed to evict safely

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

A node with low CPU use is not automatically removable. Its Pods may have large resource requests, a restricted zone, local storage, scarce attach slots, or a PodDisruptionBudget that currently allows no voluntary eviction. A consolidation controller that respects the Eviction API can wait for those constraints, while an involuntary node loss remains possible. Cost savings require a post-move node count, not just an eviction attempt.

Operational decision

A parcel API has one lightly used node beside a stateful ledger and a scheduled batch worker. Before consolidating, list requests, node affinity, topology spread, volume attachments, PDB disruptions allowed, and the surge reserve for the next release. The read-only commands below are a starting inventory; test eviction and replacement in an isolated pool. Move a replayable worker first, verify the queue lease and workload latency, and only then attempt the API Pod. If no destination node has enough allocatable resources, consolidation will fail or trigger a new node, producing no saving. Check how the provider bills partial node-hours and any minimum node-pool size. Keep capacity for a node failure and rollback after consolidation. If a PDB blocks a voluntary eviction, adjust workload availability or schedule maintenance; do not bypass it with direct Pod deletion to obtain a short-lived saving. Record exactly which node-hours were removed and which replacement resources were added.

bash
kubectl -n parcel get pdb -o wide
kubectl -n parcel get pods -o wide
kubectl get nodes -L topology.kubernetes.io/zone
kubectl -n parcel describe pod parcel-api-canary

Cost and verification

Consolidation may reduce node-hours but increase scheduling latency, hotspot risk, or recovery time. A drain also creates temporary duplicate capacity while replacements start, so the billing benefit may begin later than the controller's decision. Measure actual nodes removed, replacement starts, disruptions allowed, pending-Pod time, p95 latency, and recovery headroom. Count repeated failed attempts as control-plane and operator cost even when they do not change the bill.

Common Mistakes

  • Do not classify a node as removable from CPU utilization alone.
  • Do not bypass PDBs or local-volume constraints to force a cheaper topology.
  • Do not report savings until the billed node count actually falls.

Connected lessons

Practice and check

devops
operations
Storage details