A Pod can remain unable to start when a node has free CPU and memory but its network plugin cannot assign a Pod address. The limit may be a subnet address pool, per-node interface or address capacity, prefix availability, or a plugin-specific allocation rule. In a VPC-addressed cluster, more nodes can consume more subnet addresses and may worsen the shortage.
Pod address capacity: diagnose network allocation before adding nodes
Operational decision
A claims rollout stalls with network setup events while the autoscaler adds nodes. Inspect the affected Pod events, CNI DaemonSet health, node Pod capacity, and subnet available addresses. The sample includes a provider read-only subnet inventory; use the corresponding control for another CNI. Compare the rollout's peak old-plus-new Pods with available addresses in every target zone, not just the region total. Where prefix allocation is enabled, confirm the required contiguous block can be allocated; a scattered count of free addresses may be insufficient. Free leaked allocations only after their owner is identified. A new subnet or secondary address range requires route, network-policy, and firewall review before use. Test one new Pod on each intended zone, then verify service routing and external egress. Keep a forecast of Pod addresses per surge, autoscaling peak, and warm-pool behavior so the next deployment has a capacity gate.
kubectl describe pod -n claims claims-api-pending
kubectl get nodes -o wide
aws ec2 describe-subnets --query 'Subnets[].{Subnet:SubnetId,Zone:AvailabilityZone,Available:AvailableIpAddressCount}' --output tableCost and verification
Larger address pools and extra subnets create routing and security administration work. Warm address reservations improve Pod startup time but reduce addresses available to other workloads. Raising a node's Pod limit without confirming interface and CNI capacity can lead to more scheduled Pods stuck at network setup. Track assignment failures, per-zone free addresses, and successful ready Pods. A scheduler placement success is not the same as a usable networked Pod.
Common Mistakes
- Do not treat spare node CPU as proof of Pod IP capacity.
- Do not keep scaling nodes into a depleted subnet.
- Do not assume aggregate free addresses form an allocatable prefix.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Node autoscaling: make pending Pods schedulable before traffic rises
- Topology spread: keep replicas out of one failure domain
- Provider quota preflight: reserve capacity for rollback and recovery
- Egress policy and DNS: restrict destinations without breaking name resolution
