Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Inference retries: bound repeated work and preserve one decision

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A retry after timeout should not trigger unlimited model work or create conflicting decisions for one logical request.

Name the logical operation

Give the caller an idempotency key scoped to tenant, endpoint and a canonical request digest. A retry of the same request can retrieve the completed decision; reusing a key with a different payload must be rejected. Preserve the decision ID, model digest and policy revision with the result. Cross-region reconciliation needs this stable identity during failover, while the API contract states the expected retry behavior.

Stop duplicate in-flight work

Check a durable or appropriately shared in-flight record before starting inference. A second request with the same key may wait for the first within its deadline or receive a retry-later state; it must not launch another expensive extraction. Set a finite record lifetime longer than expected retry windows and protect the record update against races. An in-memory dictionary on one worker cannot deduplicate across replicas or regions. Region identity also determines whether replay of an old decision is allowed after policy changes.

Apply quotas at multiple levels

Bound request rate, in-flight work, queued work and cumulative compute by tenant and model. A small fast request repeated thousands of times can be as harmful as one oversized document. Reject excess before expensive stages and return a stable overload response. Do not count a cached replay as fresh model capacity, but include it in the caller rate budget if abuse is possible. Payload admission handles individual requests; queue limits keep accepted work bounded.

Verify conflict and timeout behavior

Send the same key twice concurrently, then send the same key with changed contents. Expect one logical decision for identical requests and an explicit conflict for changed contents. Force the first response to be lost after its decision was committed; retry should return the prior decision identity. The project tests that repeated receipts do not inflate model work or create a second route. A caller must know whether a timeout means “not processed” or “unknown”; it should never assume the former.

Implementation

python
from hashlib import sha256

def replay_status(records, tenant, key, payload):
    identity = (tenant, key)
    digest = sha256(payload).hexdigest()
    existing = records.get(identity)
    if existing is None:
        records[identity] = {"digest": digest, "decision": "decision-47"}
        return "new"
    if existing["digest"] != digest:
        return "conflict"
    return "replay:" + existing["decision"]

records = {}
assert replay_status(records, "merchant-82", "key-47", b"receipt-A") == "new"
assert replay_status(records, "merchant-82", "key-47", b"receipt-A") == "replay:decision-47"
assert replay_status(records, "merchant-82", "key-47", b"receipt-B") == "conflict"

Performance and operating cost

Digesting b payload bytes costs O(b) time and O(1) extra hash space; the record lookup is O(1) expected time. A durable shared record uses O(k) storage for k active keys and needs atomic creation. The example is single-process and deliberately lacks cross-worker synchronization; production deduplication must enforce one writer or transactional uniqueness.

Common Mistakes

  • Generating a new key on each retry.
  • Accepting one key with a different payload.
  • Using worker-local memory as cross-region deduplication.
  • Retaining keys for less time than clients retry.

Read next

Continue the workflow: Inference evasion response: rate limits, review and reversible containment.

Continue the workflow: Prediction API query review: distinguish probing from legitimate bursts.

ai-data
mlops
Storage details