Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Async result ledgers: reconcile output, failure and notification

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A notification is a hint to fetch a result, not the result itself; the ledger must account for every accepted job.

Write terminal state once

A worker should commit either a result reference or a failure record with job ID, input digest, model digest, policy revision and completion time. A retry may race with the first attempt, so use a conditional terminal transition or an equivalent uniqueness rule. Do not overwrite a succeeded result with a later timeout notification. The job contract defines accepted and terminal states; idempotency gives duplicate submissions one logical identity.

Treat delivery channels separately

The output object, failure object and notification may be written or delivered at different times. A success notification can arrive twice; a notification can be lost even though the result exists. Reconcile accepted jobs against durable terminal records and object existence, not notification counts alone. A missing output after a success marker is a consistency incident. The telemetry coverage ledger teaches the same denominator discipline: accepted work is the population, not observed notifications.

Replay without double action

If a job fails due to temporary input download or output upload trouble, retry it within a bounded attempt budget using the same job ID and pinned identities. A permanent invalid payload should enter a reviewable failure state. When a result feeds a downstream payment or review action, the consumer must deduplicate by decision ID; rerunning inference can produce another score but must not trigger the action twice. Decision reconciliation checks divergent routes during retries and regional recovery.

Audit aged gaps

Report counts for accepted, running, succeeded, failed, expired and canceled jobs by submission window. Flag jobs beyond expected age without terminal state, and check result-object retention before retrying. Preserve failure cause and attempt history so an operator can distinguish compute failure from lost notification. The project drops one notification and one output upload, then reconciles both without claiming the entire queue completed successfully. Include an owner and escalation time for stuck jobs.

Implementation

python
def reconcile_async_jobs(accepted, terminal, output_exists):
    gaps = []
    invalid = []
    for job_id in accepted:
        record = terminal.get(job_id)
        if record is None:
            gaps.append(job_id)
        elif record["state"] == "succeeded" and not output_exists.get(job_id, False):
            invalid.append(job_id)
    return {"missing_terminal": sorted(gaps),
            "missing_output": sorted(invalid)}

accepted = {"job-47", "job-82"}
terminal = {"job-47": {"state": "succeeded"}}
assert reconcile_async_jobs(accepted, terminal, {"job-47": True}) == {
    "missing_terminal": ["job-82"], "missing_output": []}
assert reconcile_async_jobs(accepted, terminal, {"job-47": False})[
    "missing_output"] == ["job-47"]

Performance and operating cost

Reconciling n accepted jobs is O(n) expected time and O(n) output space in the worst case; sorting gap lists adds O(g log g) for g gaps. Durable result objects and terminal records cost storage until retention expires. Scanning large queues can be partitioned by submission window, but the accepted-job denominator must remain complete.

Common Mistakes

  • Counting notifications as completed jobs.
  • Overwriting an existing success with a delayed retry failure.
  • Replaying a downstream side effect when an inference result is redelivered.
  • Deleting failure records before the reconciliation window closes.

Read next

Continue the workflow: Project: explain a receipt-model serving cost spike.

ai-data
mlops
Storage details