Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Container OOM attribution: separate a limit kill from node eviction

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

A container OOM kill occurs when memory charged to its cgroup cannot be reclaimed within its limit. A node-pressure eviction is a kubelet decision made from node availability and Pod priority. They can appear in the same incident, but they have different remedies. A restarted container can also have exited for an application error or a probe failure. The reason field alone is a starting point; capture the preceding trend and the node condition before the old container state disappears.

Operational decision

A document-conversion Pod restarts every time a 73 MiB scan is decoded. Record its container ID, previous termination reason, restart count, Pod events, node MemoryPressure condition, and memory usage by container. If the container was OOMKilled while the node stayed healthy, inspect its limit and charge breakdown; if the Pod was Evicted after node pressure, inspect aggregate requests and node reservations. Repeat one controlled load in a disposable cluster and preserve the memory peak, input size, and result. Do not erase the fault by increasing the limit until you know whether the process retained data or simply needs a documented working set.

bash
kubectl -n claims describe pod converter-0
kubectl -n claims get pod converter-0 -o jsonpath='{range .status.containerStatuses[*]}{.name}{" previous="}{.lastState.terminated.reason}{" restarts="}{.restartCount}{"\n"}{end}'
kubectl describe node worker-07

Cost and verification

The three API reads are O(1) in object count, but event retention and monitoring resolution limit historical certainty. Each retry consumes startup CPU and may replay work; correlate OOM frequency with duplicated jobs and user latency. A larger limit increases worst-case node demand and can move the failure to another Pod. Preserve a pre-change baseline and verify the same input after the fix.

Common Mistakes

  • Do not equate an exit code with a complete root cause.
  • Do not call an eviction a container-limit OOM.
  • Do not raise a limit without checking node headroom and input replay.

Connected lessons

Practice and check

devops
memory-operations
Storage details