Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: publish and replay a nightly receipt scoring partition

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Run a versioned receipt scoring batch, recover from partial failure and reconcile a corrected replay before consumers act.

Set the batch contract

At 02:00, score a frozen receipt partition with one verified model digest and feature-contract version. Give the job a key derived from partition, data, model and code identities. Each input must end as scored, rejected or held, with a stable receipt ID. Idempotent publication keeps partial output invisible until counts and uniqueness checks pass.

Prepare failure fixtures

Include 529 valid receipts, one malformed amount, duplicate IDs, a worker crash halfway through a write and a retry with the same job key. A corrected source snapshot removes one receipt and changes another. Freeze expected output counts and changed-ID sets before running. Use a distinct model digest for a comparison-only replay to prove that comparison results never move the production pointer.

Stage, verify and publish

Validate input identity and model artifact, write a temporary output, reconcile source IDs with scored and rejected IDs, then publish one immutable manifest. Retry a failed job without appending duplicate results. For a corrected replay, compare old and new outputs and require downstream consumers to acknowledge the supersession protocol. Replay reconciliation exposes removed, added and changed decisions.

Test consumer behavior

Simulate a manual-review service keyed by receipt and decision version. The same completed batch must not create a second review task. A correction may update an existing task only under an explicit rule; a comparison run must create none. Report time to complete, per-record failures, duplicate prevention, diff counts and the active output manifest. Keep the old version available for audit and rollback until retention policy permits deletion.

Implementation

python
def batch_accounting(input_ids, scored_ids, rejected_ids):
    inputs = list(input_ids)
    scored, rejected = set(scored_ids), set(rejected_ids)
    if len(inputs) != len(set(inputs)):
        return {"state": "hold", "reason": "duplicate-input"}
    if scored & rejected:
        return {"state": "hold", "reason": "conflicting-outcome"}
    if scored | rejected != set(inputs):
        return {"state": "hold", "reason": "incomplete-partition"}
    return {"state": "publish", "input_count": len(inputs)}

receipts = ["receipt-47", "receipt-82", "receipt-129"]
assert batch_accounting(receipts, receipts[:2], receipts[2:])[
    "state"] == "publish"
assert batch_accounting(receipts, receipts[:1], receipts[2:])[
    "state"] == "hold"

Performance and operating cost

Accounting uses O(n) expected time and O(n) space for n input IDs. Inference, storage and downstream reconciliation dominate a real batch. This set-based check is local; atomic publication and cross-worker idempotency require durable conditional writes or database constraints.

Common Mistakes

  • Exposing half a partition after a worker crash.
  • Dropping malformed records without a rejected outcome.
  • Creating duplicate review tasks on retry.
  • Publishing a comparison replay as if it replaced production decisions.

Read next

Continue the workflow: Project: retire a receipt model without breaking batch or audit.

Continue the workflow: Project: reconcile a queued receipt classification job.

ai-data
mlops
Storage details