A HorizontalPodAutoscaler adjusts a workload's desired replica count from observed metrics. CPU utilization compares usage with declared CPU requests, so missing or unrealistic requests distort that signal. Scaling adds capacity only after metrics arrive, the controller acts, and new Pods start and become ready.
Horizontal autoscaling: choose a signal tied to demand
Operational decision
For a thumbnail service, CPU may correlate with work when each request has similar processing cost. Measure the CPU request, per-Pod throughput, and cold-start duration before choosing a target. Keep a minimum replica count that can handle the interval before new Pods become ready. A queue worker may need queue age or backlog per worker instead of CPU, because an idle process can still face a growing backlog. The YAML sample scales a Deployment between three and eleven replicas at a 68 percent average CPU-utilization target. Verify the metrics adapter, Pod requests, and application readiness before applying it. During a load test, inspect whether latency improves as replicas rise and whether downstream storage becomes the new bottleneck. A rising HPA desired count with pending Pods is a capacity fault, not a successful scale-out.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: {name: thumbnail-api}
spec:
scaleTargetRef: {apiVersion: apps/v1, kind: Deployment, name: thumbnail-api}
minReplicas: 3
maxReplicas: 11
metrics:
- type: Resource
resource:
name: cpu
target: {type: Utilization, averageUtilization: 68}Cost and verification
More replicas consume compute and can multiply connections to a shared database. Scaling has a delay; it cannot rescue traffic spikes shorter than the measurement and startup path. Overreactive rules can oscillate, so test scale-down behavior and stabilization on the target cluster. The sample requires a functioning resource-metrics pipeline. If the service is dominated by a single serialized dependency, additional Pods can make contention worse while the autoscaler reports progress.
Common Mistakes
- Do not use CPU utilization without CPU requests.
- Do not assume desired replicas are running and ready.
- Do not scale the front end past a fixed downstream bottleneck.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Kubernetes requests and limits: schedule for real load
- Kubernetes probes: startup, readiness, and liveness
- Queue consumers: acknowledgement, idempotency, and backlog
