Kubernetes schedules a Pod against its resource request. An HPA CPU utilization target compares measured CPU use with requested CPU, so a request change can alter desired replicas even when traffic and actual CPU are unchanged. Shrinking requests to fit more Pods may increase the reported utilization and trigger more replicas; raising requests can delay scaling while reserving more node capacity. A cost comparison that ignores this control loop is incomplete.
CPU request rightsizing: account for the HPA feedback loop
Operational decision
A parcel-rating API uses a CPU-utilization HPA and sometimes saturates during carrier batch imports. Record per-Pod CPU, p95 latency, throttle time, replica count, node count, and pending Pods over a normal week and a known peak. Apply a small request change to one canary Deployment, leaving the HPA target and load shape explicit. The fragment represents the container request and limit; use it inside the reviewed workload manifest. Compare the same workload before and after the change, including new-node startup and rollback surge. If sidecars use CPU, check whether the HPA uses Pod-level or container-level utilization so the denominator matches the intended work. Do not combine request reduction with a simultaneous algorithm change; that destroys attribution. Verify memory separately, because an inexpensive CPU request cannot prevent an OOM kill. Roll back if latency or error budget deteriorates even when billed node-hours fall.
resources:
requests:
cpu: 340m
memory: 620Mi
limits:
memory: 1GiCost and verification
The apparent saving from lower requested CPU may vanish when the HPA creates more replicas or when node fragmentation leaves the same number of nodes running. Conversely, a larger request can reduce replica churn but strand capacity. Measure cost per successful rating at a fixed demand level, ready replicas at peak, scheduler Pending time, and p95 latency. The sample values are a canary hypothesis, not a sizing recommendation; derive production requests from observed load and failure headroom.
Common Mistakes
- Do not adjust CPU requests without checking the HPA's utilization denominator.
- Do not infer a smaller bill from fewer requested cores alone.
- Do not use average CPU as the only sizing input when tail latency and cold starts matter.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Kubernetes requests and limits: schedule for real load
- Horizontal autoscaling: choose a signal tied to demand
- HPA stabilization: prevent replica oscillation without masking demand
- Shared cluster cost: reconcile service allocation with the provider bill
