A HorizontalPodAutoscaler changes a workload's desired replica count from sampled metrics. Its control loop is periodic, and new Pods need scheduling, image pulls, startup, and readiness before they add usable capacity. A scale-down stabilization window holds recent recommendations to avoid rapid replica loss after a short dip. It does not reserve nodes or increase database sessions.
HPA stabilization: prevent replica oscillation without masking demand
Operational decision
A parcel API sees short traffic troughs between campaign batches. Configure a conservative scale-down window and a maximum reduction per interval, then test with a load trace that includes two bursts separated by a quiet period. The YAML fragment belongs under an autoscaler specification; it is not a complete object. Set CPU requests to measured values before using utilization as a signal, and observe unavailable endpoints and p95 latency as the controller changes desired replicas. Compare the chosen window with Pod startup and node-provisioning delays. If the autoscaler requests more Pods but they remain Pending, inspect scheduling and node pool limits rather than shortening the scale-down window. If a database pool is fixed, cap replicas or reduce per-Pod sessions so scale-out does not exhaust the backend. Validate missing-metric behavior in a disposable environment; the controller can make a different decision when some metrics are unavailable.
spec:
minReplicas: 4
maxReplicas: 19
behavior:
scaleDown:
stabilizationWindowSeconds: 360
policies:
- type: Pods
value: 2
periodSeconds: 60Cost and verification
Longer stabilization keeps extra Pods running through quiet periods, raising compute cost but reducing cold scale-out and endpoint churn. A short window saves capacity and can cause repeated initialization work or latency spikes. Evaluate the full response time from load rise to Ready endpoints, not only the autoscaler's desired count. Revisit thresholds when request sizes, initialization cost, or downstream capacity changes. Record whether the chosen metric reflects actual user demand or a symptom of an overloaded dependency.
Common Mistakes
- Do not assume desired replicas are immediately serving traffic.
- Do not tune HPA CPU utilization without realistic CPU requests.
- Do not scale API replicas beyond the database session budget.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Horizontal autoscaling: choose a signal tied to demand
- Node autoscaling: make pending Pods schedulable before traffic rises
- Database pool pressure: bound waiting before the database collapses
- Capacity and load tests: identify the next bottleneck
