On a Linux cgroup v2 host, memory.current shows charged memory, memory.stat separates broad charge classes, and memory.events reports boundary events such as high, max, and OOM. Pressure stall information records time tasks could not progress because of resource pressure. A rising RSS graph without a matching latency or reclaim signal is different from a service repeatedly stalling on reclaim. These counters are scoped to a cgroup; they must be read from the workload's actual path, not an unrelated host shell.
Cgroup memory signals: read reclaim before the OOM counter
Operational decision
A settlement worker slows during the hourly ledger sweep, yet no OOM kill occurs. In an isolated test, collect memory.current, memory.stat, memory.events, and memory.pressure before, during, and after the sweep. Calculate deltas rather than treating cumulative counters as instantaneous rates. Correlate reclaim and full-stall changes with queue age and request latency. If the pressure is caused by file cache, compare active and inactive file charge with working-set demand; if anonymous memory grows and stays high after the sweep, investigate retained objects or allocator behavior. Preserve node-level pressure separately: one container may look quiet while its node struggles.
cat /sys/fs/cgroup/memory.current
cat /sys/fs/cgroup/memory.stat
cat /sys/fs/cgroup/memory.events
cat /sys/fs/cgroup/memory.pressureCost and verification
Reading four small control files is O(1), but storing per-container samples at high frequency scales with active cgroup count and retention duration. Sample at a rate that preserves short sweeps without multiplying monitoring cost. Use workload-level latency and queue age as outcome signals; a counter increase without user impact may call for capacity planning rather than paging. Check cgroup version and metric exporter semantics before comparing clusters.
Common Mistakes
- Do not read host counters and label them as container counters.
- Do not interpret a cumulative event count without a time window.
- Do not treat OOM absence as proof that memory pressure is harmless.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Container OOM attribution: separate a limit kill from node eviction
- Observability: join metrics, logs, and traces
- Metric cardinality: keep observability usable during a surge
