A node autoscaler provisions or removes nodes in response to unschedulable Pods and configured node constraints. It cannot create capacity for a Pod whose requests, topology rules, volume limits, or node selectors have no compatible node type. Workload autoscaling and node autoscaling act on different signals and time scales; more desired replicas do not mean immediately available service capacity.
Node autoscaling: make pending Pods schedulable before traffic rises
Operational decision
A parcel API scales from six to fifteen replicas during a promotion. Model the peak request footprint, including one surge replica per rollout slice, then check that at least one permitted node type can fit each pending Pod. The command sequence is a diagnostic read, not an autoscaler installation. Inspect scheduling events for the first Pending Pod, node-pool upper bounds, topology spread, and storage attachment limits. Measure the time from a Pending Pod to a Ready endpoint under a controlled load ramp. If provisioning takes several minutes, keep minimum spare capacity or pre-scale before a known event. Confirm that the ingress and database can support the added replicas; otherwise the autoscaler only moves the bottleneck. Test scale-down separately so draining does not violate a stateful application's availability rule.
kubectl get pods -n parcels --field-selector=status.phase=Pending
kubectl describe pod -n parcels parcel-api-pending-47
kubectl get nodes -L topology.kubernetes.io/zone
kubectl get events -n parcels --sort-by=.lastTimestampCost and verification
Spare nodes cost money while idle, but cold provisioning adds latency and may fail during provider capacity shortages. Under-requesting resources makes scaling appear cheaper until contention or eviction hits. Over-requesting can make otherwise healthy Pods unschedulable. Record scale-out lead time and compare it with the service's tolerated demand spike. Readiness after scheduling still depends on image pulls, startup, and downstream availability.
Common Mistakes
- Do not assume a Pending Pod will trigger a node that can satisfy impossible affinity.
- Do not treat desired replicas as available endpoints.
- Do not tune only CPU scaling while database sessions remain fixed.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Horizontal autoscaling: choose a signal tied to demand
- Node pressure eviction: trace lost Pods to exhausted local resources
- Namespace quotas: reserve room for a safe rollout
- Database pool pressure: bound waiting before the database collapses
Practice and check
Advanced follow-up
Advanced follow-up
Advanced follow-up
Advanced follow-up
- Interruptible compute: price recovery work, not only cheap node hours
- Node consolidation: calculate the capacity needed to evict safely
