Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Node pressure eviction: trace lost Pods to exhausted local resources

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

Node-pressure eviction is a kubelet response to scarce node resources such as available memory or filesystem space. The kubelet may reclaim local resources and then terminate Pods when configured thresholds remain crossed. Pod requests, actual use, and priority affect which Pods are candidates. A PodDisruptionBudget does not prevent this involuntary termination, so an apparently safe maintenance budget is not a node-capacity plan.

Operational decision

A document-index worker disappears during a large image rollout. Inspect Pod termination reason, node conditions, events, and filesystem usage before restarting it. The commands are read-only; run them in the affected cluster context. If ephemeral storage is full, find whether image layers, writable container data, or logs consumed it. Set realistic ephemeral-storage requests and limits, and give the node room to pull the next image before draining another node. If memory pressure caused the eviction, distinguish container OOM from node-level eviction and check whether request values reflect observed working sets. Moving the worker to another full node only shifts the incident. Verify replacement Pods process a synthetic document and that queue age falls after capacity is restored.

bash
kubectl get pod -n indexing -o wide
kubectl describe pod -n indexing document-index-worker-0
kubectl describe node worker-node-47
kubectl get events -n indexing --sort-by=.lastTimestamp

Cost and verification

Keeping free node memory and disk has a direct compute cost, but a zero-headroom cluster pays through failed image pulls, evictions, and slower recovery. Tight limits can terminate healthy bursts; missing requests invite unsafe packing. Track node resource pressure and Pod eviction counts with the user-facing queue age. Do not infer an application regression solely from a restart count, and do not call an eviction solved until node pressure and replacement capacity are both understood.

Common Mistakes

  • Do not expect a disruption budget to block node-pressure eviction.
  • Do not treat every terminated Pod as an application OOM.
  • Do not ignore image and log disk use when sizing ephemeral storage.

Connected lessons

Practice and check

Advanced follow-up

Memory failure follow-up

devops
operations
Storage details