A cloud control plane can reject bursts even when the requested infrastructure has room to exist. A 429 or provider throttling response is a rate signal; a quota-exceeded response may require a limit increase or design change. Concurrent deploy jobs, inventory scans, and autoscalers can share an account or regional request budget. Blind retries can turn one slow rollout into a control-plane flood.
Cloud API throttling: keep infrastructure changes inside a request budget
Operational decision
A regional failover job creates instances while several pipelines refresh infrastructure state. Record the operation, account, region, response code, request ID, and retry count without storing credentials. Cap concurrent mutating calls per operation family. Where supported, honor server retry guidance; otherwise use bounded exponential delay with jitter and a total deadline. The text policy is an operational contract, not a provider API configuration. Read back a timed-out create before resending it, and retain the same idempotency token if the API supports one. Pause nonessential inventory scans while preserving health and rollback operations. Test that one throttled job eventually stops and hands an actionable record to an operator instead of retrying forever. A successful retry proves that request succeeded; it does not prove the whole rollout reached its intended state.
Control-plane request policy
Scope: account + region + operation family
Concurrent mutations: 4 per scope
Retry: transient throttle only; jittered delay; 90-second total deadline
Ambiguous create: read back by stable operation key before retry
Failure record: request ID, last response, resources found, next safe actionCost and verification
Lower concurrency may slow provisioning but leaves capacity for rollback and incident response. Extra retries consume the same quota they are trying to survive. Rate budgets must include SDK retries, orchestration retries, and human-run tools, or the effective attempt count can multiply. Do not place every independent service behind one adaptive limiter unless that shared behavior is intended. Measure request rate, throttled responses, rollout duration, and desired-versus-actual resources together.
Common Mistakes
- Do not retry a hard service quota error as though it were transient throttling.
- Do not add an outer retry loop without counting SDK attempts.
- Do not resend an ambiguous create with a fresh operation identity.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Terraform state: shared ownership and safe plans
- Retries and timeouts: bound the cost of a failed request
- Node autoscaling: make pending Pods schedulable before traffic rises
- Multi-region failover: define write ownership before moving traffic
