Skip to content
AITroveRead. Build. Understand.
Make this comfortable

CPU request rightsizing: account for the HPA feedback loop

Last updated: 1 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

Kubernetes schedules a Pod against its resource request. An HPA CPU utilization target compares measured CPU use with requested CPU, so a request change can alter desired replicas even when traffic and actual CPU are unchanged. Shrinking requests to fit more Pods may increase the reported utilization and trigger more replicas; raising requests can delay scaling while reserving more node capacity. A cost comparison that ignores this control loop is incomplete.

Operational decision

A parcel-rating API uses a CPU-utilization HPA and sometimes saturates during carrier batch imports. Record per-Pod CPU, p95 latency, throttle time, replica count, node count, and pending Pods over a normal week and a known peak. Apply a small request change to one canary Deployment, leaving the HPA target and load shape explicit. The fragment represents the container request and limit; use it inside the reviewed workload manifest. Compare the same workload before and after the change, including new-node startup and rollback surge. If sidecars use CPU, check whether the HPA uses Pod-level or container-level utilization so the denominator matches the intended work. Do not combine request reduction with a simultaneous algorithm change; that destroys attribution. Verify memory separately, because an inexpensive CPU request cannot prevent an OOM kill. Roll back if latency or error budget deteriorates even when billed node-hours fall.

yaml
resources:
  requests:
    cpu: 340m
    memory: 620Mi
  limits:
    memory: 1Gi

Cost and verification

The apparent saving from lower requested CPU may vanish when the HPA creates more replicas or when node fragmentation leaves the same number of nodes running. Conversely, a larger request can reduce replica churn but strand capacity. Measure cost per successful rating at a fixed demand level, ready replicas at peak, scheduler Pending time, and p95 latency. The sample values are a canary hypothesis, not a sizing recommendation; derive production requests from observed load and failure headroom.

Common Mistakes

  • Do not adjust CPU requests without checking the HPA's utilization denominator.
  • Do not infer a smaller bill from fewer requested cores alone.
  • Do not use average CPU as the only sizing input when tail latency and cold starts matter.

Connected lessons

Practice and check

devops
operations
Storage details