Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: reconcile a queued receipt classification job

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Submit slow receipt scans as immutable jobs, inject output and notification failures, and account for every accepted request.

Define the job envelope

A bulk receipt scan can take longer than an interactive deadline. Accept a versioned input object and return a job ID, submission time and status location. Pin the image digest, extraction chain revision, risk model digest and policy revision at submission. Set a queue-age limit tied to the reconciliation cutoff. The job contract makes accepted distinct from completed; chain identity prevents a worker from mixing stage versions later.

Inject three failures

Submit three jobs. Allow one to succeed and notify normally; complete one but drop its notification; let one model invocation finish while its output upload fails. Also queue a fourth until it expires. Retry the failing job with the same pinned identities and bounded attempt count. A consumer must poll or reconcile durable state rather than wait forever for a notification. The ledger identifies missing terminal records and a success marker without an output object.

Protect the downstream action

The review system consumes completed results by decision ID. Deliver the same success twice and verify it creates one review task. If the output changes between attempts under the same job ID, quarantine the conflict rather than silently choosing the latest score. Retry control prevents repeated expensive work at ingress, while decision reconciliation prevents repeated or divergent actions after result delivery.

Publish job accounting

Report accepted, terminal, expired, failed, retried and unresolved counts; input and model digests; output-object checks; notification delivery; consumer deduplication and time to completion. If a job has no terminal record, state that explicitly rather than treating it as failed or succeeded. The project passes only when every accepted job is terminal or has an owned open investigation. Link capacity and queue age to operator alerts before enabling a larger submitter.

Implementation

python
def async_release_report(accepted, succeeded, failed, expired):
    terminal = succeeded | failed | expired
    if (terminal - accepted or (succeeded & failed) or (succeeded & expired)
            or (failed & expired)):
        return {"state": "invalid", "unresolved": []}
    unresolved = sorted(accepted - terminal)
    return {"state": "investigate" if unresolved else "accounted",
            "unresolved": unresolved}

accepted = {"job-47", "job-82", "job-93"}
assert async_release_report(accepted, {"job-47"}, {"job-82"},
                            {"job-93"})["state"] == "accounted"
assert async_release_report(accepted, {"job-47"}, set(),
                            {"job-93"})["unresolved"] == ["job-82"]

Performance and operating cost

Set operations take O(n) expected time and O(n) memory for n jobs. The real cost is input storage, queued worker capacity, retry compute, result storage and notification handling. Scale-out can improve throughput but does not make result commits atomic; verify durable state transitions and consumer idempotency separately.

Common Mistakes

  • Assuming a success notification proves the output object exists.
  • Letting queued jobs use whatever model alias is current at execution time.
  • Counting duplicate success deliveries as separate reviews.
  • Hiding a job with no terminal record inside a generic failure total.

Read next

ai-data
mlops
Storage details