Use a disposable Kubernetes cluster and a document-conversion workload with a memory-backed scratch volume. The goal is to identify three different failure boundaries without hiding one by moving it to another Pod. Save the workload manifest, input sizes, node allocatable capacity, and monitoring snapshots before changing any limit.
Project: diagnose a memory failure without moving it to another Pod
Inject and classify
First, use a bounded input that crosses the container memory limit; capture prior termination state, cgroup counters, restart count, and node condition. Next, leave task files in tmpfs across a container restart and show that file charge remains while the heap recovers. Finally, schedule a two-Pod surge against a node pool that cannot fit both requests. Explain which fault is a cgroup kill, which is scratch growth, and which is an admission or placement failure. Keep a separate node-pressure run if the test cluster can safely support it.
Memory drill acceptance
Input: fixed 73 MiB scan plus a smaller control input
Kill: previous container state and charge peak recorded
Scratch: abandoned file persists across container restart
Placement: requested surge exceeds per-node fit
Recovery: cleanup and capacity change tied to each cause
User path: one accepted conversion after every repairRepair and recheck
Set a scratch size limit and task cleanup, then replay the failed input and verify that a full volume produces a handled error. Budget JVM heap and native headroom from observed peaks if the renderer uses Java. For the rollout, add capacity or reduce surge only after testing the service availability target. Run the same inputs three times and report post-cycle memory, pressure stall change, OOM counts, queue age, and p99 latency. A repair passes only when it removes the targeted failure and keeps the user path working without raising another Pod's failure rate.
Common Mistakes
- Do not claim a restart fixed a retained-memory problem.
- Do not use one node-wide memory graph to attribute a container kill.
- Do not reduce requests to make a rollout appear schedulable.
Connected lessons
- Container OOM attribution: separate a limit kill from node eviction
- JVM container memory: leave room beyond the Java heap
- Memory-backed emptyDir: budget file bytes as container memory
- Cgroup memory signals: read reclaim before the OOM counter
- Memory growth: distinguish retained data from useful cache
- Memory-safe rollouts: reserve space for old and new Pods
- DevOps projects
