Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Provisioning state machines: handle partial success and retries

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

A template runner can time out after creating a repository but before creating its database, or after issuing a provider request but before receiving the response. Blindly rerunning the whole script can duplicate resources or erase the only clue to an in-flight operation. Model each step with a durable operation ID, desired identity, observed provider identity, retry policy, timeout, and compensating action. Compensation is not always deletion: a database holding customer data may need retention and human review after a later step fails.

Operational decision

For an audit-reporting service, the workflow creates a repository, workload identity, namespace, database, and deployment binding. Persist step state before and after each external call. Force the database response to time out after the provider has accepted it, then resume by looking up the stable requested database identity rather than issuing another create. Force the deployment binding to fail after the database exists. Mark the request partially ready, keep the database under a bounded retention rule, and present a repair or disposal choice to the owner. The workflow should never report Ready until a synthetic application connection succeeds and the owner can locate the created resources. A cancellation request must stop new steps while allowing in-flight calls to settle and be inventoried. Record every resource's desired identity, observed ID, creation timestamp, cost owner, and final disposition; otherwise later cleanup becomes guesswork. Reconcile the record against the provider inventory periodically to detect resources that exist but were never recorded as complete.

Output
Audit service states
Requested -> Validated -> Provisioning
Provisioning -> Ready | PartialFailure | CancelPending
PartialFailure -> Repairing | RetainedForReview
CancelPending -> Settled -> Cleaned | RetainedForReview
Every transition: request ID, step ID, provider ID, timestamp
Ready gate: resource inventory plus application connection

Cost and verification

A workflow with K steps needs O(K) durable state and provider calls; retries increase requests and can amplify cost if external APIs lack stable identifiers. Reconciliation work scales with outstanding requests plus known resources, so archive terminal records without losing audit evidence. Measure stuck transitions, duplicate provider IDs per request, partial failures older than their repair window, and Ready states whose user-path probe fails. A timeout is uncertainty, not proof that a provider did nothing.

Common Mistakes

  • Do not rerun a timed-out create without checking the provider for the intended resource.
  • Do not delete stateful resources as a blanket compensation action.
  • Do not equate the final script exit code with the actual resource inventory.

Connected lessons

Practice and check

devops
platform-engineering
Storage details