A network timeout after a cloud create request does not establish whether the provider accepted the mutation. Repeating the call with a new identity can allocate a second resource, incur cost, and leave automation state inconsistent. The recovery unit is the desired resource identity plus an idempotency token or stable tag set, not the HTTP response alone.
Ambiguous cloud creates: reconcile before repeating a timed-out mutation
Operational decision
A deployment tool creates a recovery load balancer and loses the response. Preserve the request token, configuration fingerprint, operation time, and provider request ID. Read the provider's resource inventory using the stable token or tags and inspect any asynchronous operation status. If exactly one matching resource exists, adopt its actual identifier into the state record and verify listeners, network attachment, and health before proceeding. If none exists after a bounded consistency window, retry with the same idempotency identity when the API supports it. If two matches exist or ownership is uncertain, stop and require a reviewed reconciliation; deleting either one automatically could remove live traffic. The text block specifies the decision tree without pretending every provider exposes identical idempotency fields. Run this as a fault-injection test by dropping the create response after the provider accepts the call. A successful local command exit is not the acceptance criterion; the inventory must contain one intended resource and the automation state must point to it.
Create reconciliation decision
Input: desired resource key, request token, configuration digest
Read: provider operation status and tagged inventory
One match: verify properties, adopt actual ID, continue
No match after consistency window: retry with same token
Multiple or mismatched matches: stop for reviewed reconciliation
Acceptance: one intended resource, one state record, healthy routeCost and verification
Read-back calls and a bounded wait add deployment time and control-plane API usage. They are cheaper than orphaned resources or duplicate routing targets. Tags alone may be nonunique or delayed, so pair them with provider operation identifiers and exact properties. Keep reconciliation idempotent across process restarts and expose unresolved mutations in the release record. Do not run cleanup by broad name prefix during an active incident.
Common Mistakes
- Do not infer a create failed solely because its response timed out.
- Do not use a fresh request token for an ambiguous retry when the API supports idempotency.
- Do not delete a suspected duplicate before checking traffic ownership.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Cloud API throttling: keep infrastructure changes inside a request budget
- Terraform state: shared ownership and safe plans
- Release evidence: tie one deployed digest to one approval decision
- Multi-region failover: define write ownership before moving traffic
