Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Asynchronous inference: bind a queued job to immutable inputs

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A queued prediction needs a stable job ID, input digest, model identity and expiry policy because acceptance is not completion.

Separate accepted from completed

A receipt archive may submit a scan that takes too long for a synchronous endpoint. The submission response acknowledges durable acceptance and returns a job ID; it is not a prediction. Define states such as accepted, running, succeeded, failed, expired and canceled, with one terminal state per logical job. Store caller, input object digest, approved model digest, feature contract, policy revision and submission time. The API contract must distinguish these states so a caller never treats an accepted job as a clear receipt.

Freeze what the worker will read

A queue can hold work while model pointers and input objects change. Pin immutable input bytes or a versioned object plus digest, and pin the model and policy approved for that job. Reject an input that has been replaced or deleted before processing; do not silently read a newer file under the old job ID. Artifact identity protects the model package, while feature admission checks the delayed input when it is actually processed. The result should state the identities it used.

Set expiry from business value

A receipt classification needed before daily reconciliation can become useless after the cutoff. Set queue age, processing timeout and result retention separately. Expire a job that waited too long instead of consuming compute and delivering a stale action. Caller retries should reuse a logical key and either return the existing job or a clear conflict. Retry identity avoids duplicate work, but asynchronous jobs also need a terminal result and a retention clock.

Test lifecycle edges

Stop the worker after claiming a job, delay the queue beyond expiry and change the current production model alias. The retried job must either complete under its pinned digest or fail with an explicit identity error, never switch models unnoticed. The result ledger reconciles terminal states and late notifications. The applied project exercises a job that expires while a newer model is already serving synchronous traffic.

Implementation

python
def async_job_state(job, now_seconds):
    if now_seconds < job["submitted_at"]:
        raise ValueError("clock before submission")
    if job["state"] in {"succeeded", "failed", "expired", "canceled"}:
        return job["state"]
    if now_seconds - job["submitted_at"] > job["max_queue_age"]:
        return "expired"
    if not job["input_digest"] or not job["model_digest"]:
        return "failed"
    return job["state"]

job = {"state": "accepted", "submitted_at": 1000, "max_queue_age": 47,
       "input_digest": "input-82", "model_digest": "risk-47"}
assert async_job_state(job, 1020) == "accepted"
assert async_job_state(job, 1048) == "expired"

Performance and operating cost

The state check is O(1) time and space. Queue storage scales with pending jobs and immutable input bytes; workers pay for warmup, inference and result writes. A stale job may be cheaper to expire than to run, but expiry must be visible to callers and counted in operating metrics.

Common Mistakes

  • Treating job acceptance as a completed model decision.
  • Resolving a mutable production alias only when a delayed worker starts.
  • Reading a changed input object under the old job ID.
  • Using one timeout for queue age, execution and result retention.

Read next

ai-data
mlops
Storage details