Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Failure classification and retry budgets

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A retry policy repeats only failures likely to clear without changing the data or code, while limiting attempts so an outage cannot flood an upstream system.

Classify before repeating

A connection reset or temporary rate limit may clear. A malformed schema, missing required field or violated uniqueness rule will fail again on the same input. Put those records in a quarantine path with context and stop the interval before publication. Quality gates handle data defects; retries handle transient execution defects.

Bound attempts and total delay

Five rapid retries against an overloaded source can amplify its failure. Use a finite attempt count, exponential delay with jitter, a deadline tied to the consumer SLO, and a circuit breaker when the upstream dependency is broadly unavailable. The deadline matters: a retry after the finance cutoff may be useful for repair, but cannot make that interval on time.

Make side effects safe

A task that posts a billing export can succeed remotely and lose its acknowledgement locally. Blind retry can send the export twice. Use a stable idempotency key understood by the receiver, or write to a staging resource and commit once. If the external endpoint cannot deduplicate, reconcile its state before repeating. Task interval identity supplies a stable key.

Separate pause from discard

When a rate limit lasts longer than the retry budget, mark the interval blocked or failed with its exact input snapshot and resumable checkpoint. Do not skip source positions to keep a dashboard green. Checkpoint discipline makes recovery possible after the dependency returns.

Exercise an uncertain outcome

In a release test, simulate the receiver accepting a request while the client times out. Re-run with the same idempotency key and confirm there is one export. Also test a permanent schema error and confirm it is not retried. Record attempt count, last error class and next action in the run log so an operator does not infer state from silence.

Implementation

python
def retry_plan(error_kind, attempt, max_attempts=4):
    permanent = {"invalid_schema", "duplicate_business_key", "missing_field"}
    transient = {"timeout", "connection_reset", "rate_limited"}
    if error_kind in permanent:
        return {"action": "quarantine", "delay_seconds": 0}
    if error_kind not in transient or attempt >= max_attempts:
        return {"action": "stop", "delay_seconds": 0}
    return {"action": "retry", "delay_seconds": min(6 * (2 ** (attempt - 1)), 48)}

assert retry_plan("timeout", 2) == {"action": "retry", "delay_seconds": 12}
assert retry_plan("invalid_schema", 1)["action"] == "quarantine"
assert retry_plan("rate_limited", 4)["action"] == "stop"

Performance and operating cost

A policy decision is O(1); each retry repeats the task's I/O and compute cost. This deterministic delay example omits jitter so its assertion is reproducible; production schedules should add bounded jitter and honor an upstream retry-after value when available.

Common Mistakes

  • Do not retry permanent data or schema errors.
  • Do not assume a client timeout means the remote side did nothing.
  • Do not use unbounded retries that outlive the consumer deadline.

Read next

Continue the workflow: Telemetry correlation and cardinality budgets.

ai-data
data-engineering
Storage details