A chain should make one defensible decision when an early stage fails, a later stage times out or the total deadline expires.
Test partial failure and deadline exhaustion in model chains
Carry remaining time forward
Start an absolute request deadline at ingress and recompute time remaining before each stage. The extraction stage can consume most of the budget on a hard scan; the risk model must not then begin a fresh unrestricted wait. Include queue time and transit overhead. A deadline exhausted before scoring should return a named fallback state without an invented score. Stage contracts establish which failures are permitted, and overload policy says how callers see the result.
Exercise the failure matrix
For each stage, test malformed input, unavailable dependency, slow response, stale artifact and wrong schema. Confirm that downstream stages are skipped when their preconditions fail and that the decision log records the first causal failure. A later timeout should not overwrite an earlier contract mismatch. Keep a control case that traverses the whole chain. Deterministic fixtures make the branches repeatable; a live canary checks deployed wiring and real resource limits.
Name the degraded decision
Manual review may be safe for an unreadable receipt but unsafe for a request that must be rejected on authorization failure. The fallback table needs owner, expected volume and test evidence. Separate “no model score was produced” from “model score below threshold” in the response and event stream. Reviewer actions later refer to what the system actually proposed. A degraded path should have its own usage counter so it cannot quietly become the normal path.
Detect correlated dependency loss
Several stages may use one feature store or authentication service. A single dependency outage can drive every stage timeout and cause the fallback queue to saturate. Test under a joint failure rather than only one isolated stage at a time. The applied chain project injects both a slow extractor and missing merchant feature, then checks deadline, fallback load and complete decision identity. The operator record should identify the first broken dependency, not merely the final HTTP status.
Implementation
def run_stage_budget(deadline_ms, elapsed_ms, planned_stage_ms):
if min(deadline_ms, elapsed_ms, planned_stage_ms) < 0:
raise ValueError("negative time")
remaining = deadline_ms - elapsed_ms
if remaining <= 0:
return {"state": "fallback", "reason": "deadline-exhausted"}
if planned_stage_ms > remaining:
return {"state": "fallback", "reason": "stage-cannot-fit"}
return {"state": "run", "budget_ms": remaining}
assert run_stage_budget(180, 140, 30) == {"state": "run", "budget_ms": 40}
assert run_stage_budget(180, 140, 47)["state"] == "fallback"
assert run_stage_budget(180, 181, 1)["reason"] == "deadline-exhausted"
Performance and operating cost
Each stage budget check is O(1) time and space. The production clock should be monotonic, and measured stage duration can differ from the plan; always enforce the actual absolute deadline. Exercising s stages over f failure classes needs O(sf) cases before combined outages, while complete tracing adds storage proportional to requests and stages.
Common Mistakes
- Resetting the full request deadline at each stage.
- Calling manual review equivalent to a low model score.
- Testing only isolated failures when stages share a dependency.
- Reporting only final status and losing the first failure cause.
Read next
- Multi-stage inference: pin each stage and its contract
- Project: verify a receipt extraction-to-decision chain
- Serving overload: bound queues and choose a fallback before time runs out
- ML tests: separate code, data, model and service failures
- Project: operate receipt scoring with a deadline and overload path
Continue the workflow: Cascade release gates: branch quality, deadlines and cost.
Continue the workflow: Generative evaluation gates: grounded claims, schemas and abstention.
