Skip to content
AITroveRead. Build. Understand.
Make this comfortable

CPU throttling: distinguish a quota ceiling from node contention

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

A container CPU limit is enforced by restricting how much CPU time its cgroup can consume in a scheduling period. A process can be throttled while the node still has spare cores and while average CPU usage looks modest. CPU requests affect scheduling and relative shares; limits cap bursts. Long request latency caused by a tight CPU quota can be mistaken for network or database delay.

Operational decision

A claims API performs signature verification during a batch of uploads. Its p95 latency rises, but database time stays flat. In a Linux container using cgroup v2, read cpu.stat and compare changes in nr_throttled and throttled_usec over a measured interval. The shell commands are read-only; the cgroup path may differ on another runtime. Correlate the counter delta with request rate, CPU usage, and user latency rather than declaring every throttled period an incident. In a disposable load test, raise the limit while keeping the request budget and replica count fixed, then check whether latency falls. If the node itself is saturated, a higher limit may steal CPU from neighbors. If the workload is bursty, a realistic request and a tested limit can preserve headroom without forcing every burst into a hard ceiling. Record the node type and architecture so a CPU comparison uses comparable hardware.

bash
cat /sys/fs/cgroup/cpu.stat
kubectl top pod -n claims -l app=claims-api
kubectl get pod -n claims -l app=claims-api -o wide

Cost and verification

More CPU capacity or looser limits can raise compute cost and neighbor contention. More replicas may help independent requests but will not remove a per-request CPU bottleneck and can exhaust database sessions. A throttling counter alone gives no user impact; inspect its rate next to latency and error signals. Repeat the same load profile after a change. If the cgroup uses another version or hierarchy, use the corresponding runtime counters rather than assuming this file exists.

Common Mistakes

  • Do not infer spare container CPU from spare node CPU when a quota is active.
  • Do not increase replicas before checking the database connection budget.
  • Do not treat one nonzero throttle counter as proof of current user impact.

Connected lessons

Practice and check

devops
operations
Storage details