Choose the conclusion supported by measured work, provider responses, and reconciled resource state. A configured control is not the same as a completed recovery.
Review these lessons
- NAT port pressure: find the shared outbound ceiling
- Cloud API throttling: keep infrastructure changes inside a request budget
- Provider quota preflight: reserve capacity for rollback and recovery
- Queue-age scaling: target completion time, not only queue length
- Retry amplification: assign one owner for each failed operation
- Circuit breaker recovery: probe capacity without reopening a flood
- Dependency bulkheads: stop one slow path consuming every worker
- Ambiguous cloud creates: reconcile before repeating a timed-out mutation
Other checks
Common Mistakes
- Do not mistake a cached fallback for a live dependency response.
- Do not retry a cloud create before checking whether it succeeded.
