Classify the memory failure before selecting a repair. Each answer should account for the container, the node, and the user path affected by the change.
Review these lessons
- Container OOM attribution: separate a limit kill from node eviction
- JVM container memory: leave room beyond the Java heap
- Memory-backed emptyDir: budget file bytes as container memory
- Cgroup memory signals: read reclaim before the OOM counter
- Memory growth: distinguish retained data from useful cache
- Memory-safe rollouts: reserve space for old and new Pods
Other checks
Common Mistakes
- Do not treat every memory graph as the same signal.
- Do not confuse a fix for one Pod with recovered service capacity.
