A batch scoring retry should either reproduce the same partition output or make a new version explicit, never append duplicate decisions.
Batch inference: make partitions idempotent and outputs identifiable
Identify a scoring job completely
A nightly receipt job needs a logical partition, input snapshot digest, model digest, feature-contract version and scoring-code revision. A calendar date alone is not enough: a corrected source may be replayed for the same date. Derive a job key from the full identity and write outputs under that key. Each prediction keeps its source record ID and model identity. Artifact verification ensures the model bytes are the intended version before a partition begins.
Publish atomically after reconciliation
Write to a temporary location, count inputs and outputs, check unique record IDs and validate failures. Publish a manifest pointer only when the partition is complete. A retry with identical identity can reuse the existing completed output; a retry after code or data changes creates a distinct version. Do not append to a shared “today” file, which can double predictions after a worker crash. A durable store should enforce uniqueness or conditional publication, not rely on a process-local set.
Make record failures visible
One malformed receipt should have an explicit rejected state rather than silently disappearing. Define whether partial success is allowed and how consumers distinguish completed, incomplete and held partitions. If an output must preserve input order, say so; otherwise join by stable record ID rather than row position. Admission rules decide which records can be scored, while batch reconciliation accounts for every input.
Drill retries and duplicate inputs
Kill a worker after it writes half its output, then rerun the same job key. Confirm only one completed manifest is visible. Feed duplicate receipt IDs and a changed model digest into the same date partition; the former must hold or deduplicate under a declared rule, and the latter must create a new output version. The nightly batch project closes the loop with downstream consumption checks.
Implementation
import hashlib
import json
def batch_job_key(partition, input_digest, model_digest, code_digest):
identity = {"partition": partition, "input": input_digest,
"model": model_digest, "code": code_digest}
if not all(identity.values()):
raise ValueError("complete batch identity is required")
encoded = json.dumps(identity, sort_keys=True).encode()
return hashlib.sha256(encoded).hexdigest()
job = batch_job_key("2026-10-06", "data-47", "model-82", "code-5")
assert job == batch_job_key("2026-10-06", "data-47", "model-82", "code-5")
assert job != batch_job_key("2026-10-06", "data-47", "model-83", "code-5")
Performance and operating cost
Building the key is O(B) in encoded identity bytes and O(B) temporary space. Scoring n records costs model-dependent inference plus O(n) reconciliation. A production idempotency guarantee needs atomic publish or a store constraint; this function creates identity but cannot prevent two workers from racing to publish.
Common Mistakes
- Using a date as the only job identity.
- Appending retry output to a shared file.
- Joining predictions back to inputs by row order.
- Treating a partition as complete without accounting for rejected records.
Read next
- Batch replay: supersede outputs without duplicating downstream actions
- Project: publish and replay a nightly receipt scoring partition
- Model artifacts: verify digest, origin and loading format
- Feature contracts: admit only usable inference records
- Prediction-outcome joins: evaluate only mature, matched decisions
Continue the workflow: Asynchronous inference: bind a queued job to immutable inputs.
Continue the workflow: Stream scoring replay: checkpoints, duplicates and output reconciliation.
