Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Poison Job Quarantine and Replay Control

Last updated: 5 Oct 20267 min read
tutorial
IntermediateBy AITrove Editorial

A poison job is work that repeatedly fails without an automatic path to success. It may contain invalid input, an obsolete schema, a missing permission, or a deterministic bug. Retrying it without limit consumes worker capacity and hides newer work behind it. Quarantine removes the job from the normal retry loop, preserves a minimal failure record, and requires a deliberate repair or replay decision. A dead-letter queue is a holding mechanism, not a data-loss policy or automatic fix.

Working case

A permit report job from release 27 reaches a release-29 worker with an unsupported field shape. Every attempt throws before rendering. The queue keeps redelivering it, and a backlog of valid reports grows. After a bounded number of attempts, the worker records schema version, job ID, failure class, and tenant scope in quarantine. It does not copy the full private case into an operator dashboard. A developer ships a reader for the old shape. An operator can now replay the same job identity after checking current authorization and retention, rather than creating an unrelated duplicate request.

Implementation boundary

javascript
function jobDisposition(failure) {
  if (failure.permanent || failure.attempts >= 3) return 'quarantine';
  return 'retry-later';
}
console.log(jobDisposition({ permanent: false, attempts: 3 }));
// Output: quarantine

Classify errors before counting them. A missing required field that cannot change is permanent; a temporary storage outage should be delayed and retried. Keep retry limits and delay schedules per failure class, not as one global number. Preserve the original job identity and effect ledger when replaying. The operator action records who approved replay, the code version, why it is safe, and the new attempt number. Guard against the job being both in quarantine and active at once. Expire quarantined private payload according to the retention policy; if data is gone, report an unrecoverable result instead of pretending to replay.

Cost and boundaries

A poison item can consume O(attempt count times work before failure); moving it out after a bounded threshold protects queue capacity. Quarantine storage grows with unresolved jobs, so monitor age and count by failure class. Replaying a batch of old jobs may create a second load spike and provider expense. Use a rate-limited replay queue and test one item before widening. Minimal diagnostics may require a secure path to inspect the underlying record when needed, but logging entire private payloads increases exposure.

Failure trace

Feed an unsupported schema version and show it becomes quarantined after the chosen attempt bound. Restore a temporary storage dependency and show that a transient job retries successfully without quarantine. Attempt to replay a quarantined job after its tenant loses permission; the worker must refuse. Replay the same item twice concurrently and prove the effect ledger still yields one result. Delete its underlying record under retention rules, then request replay and return an honest unrecoverable state. Verify valid jobs continue while the poison item is isolated.

Verification

  • Permanent failures do not block healthy work indefinitely.
  • Replay retains job and effect identity.
  • Operators see bounded diagnostics and an audit trail.

Practice drill

Create 47 report jobs, one with an old schema version. Limit that class to three automatic attempts. Measure queue age for the other 46 jobs. Add a compatibility reader, authorize one replay, and record the outcome without private payload in the operator log. Repeat the replay action and verify it cannot send a second completion notice. Define when a quarantined record is deleted and what the reviewer sees afterward.

Decision note

Quarantine is a controlled exception workflow with retention and replay authority, not a hidden error bucket.

Common Mistakes

  • Retrying invalid input without a stop condition.
  • Copying private payloads into a broad operator queue.
  • Replaying with a new idempotency identity.

Related lessons

Background Workflow Reliability; Job Admission, Idempotency, and Status Resources; Worker Leases, Retries, and Duplicate Effects; Scheduled Job Overlap, Cancellation, and Compensation; Telemetry Shapes, Redaction, and Cardinality; Data Retention, Export, and Erasure.

Connected practice

Build Project: permit report job recovery and review Web Development: jobs and abuse decisions quiz.

web-tech
web-development
Storage details