Pin extraction, normalization, risk scoring and route policy as one release, then inject stage failures and measure the complete deadline.
Project: verify a receipt extraction-to-decision chain
Freeze a chain manifest
Record the extractor digest, normalization revision, risk model digest, feature contract, route policy revision and endpoint contract. Prepare a frozen set of clean scans, unreadable scans, missing merchant fields and unsupported currencies. Each fixture needs an expected route and reason, not only an expected numeric score. The chain contract binds intermediate outputs to the exact downstream consumer. Keep the previous complete manifest as the rollback target.
Break one stage at a time
Make extraction return unreadable; scoring must not run. Make normalization return the wrong schema; the chain must hold rather than coerce a value. Delay merchant lookup until only 18 milliseconds remain in the endpoint budget; the later risk stage must not start with a fresh timeout. Record the first failure, elapsed time and fallback state. The failure matrix names every planned branch before the test is run.
Test combined failure and canary
During a small live canary, pause the shared feature store while sending a hard scan. Count joint failures and manual-review load. Confirm that the public response says a degraded route was used and that logs retain all stage digests. If fallback saturates review staff, stop the canary. The service-level exercise provides the latency target, and decision logging gives a restricted trace across stages.
Choose rollout or hold
Deliver the manifest, fixture results, observed stage latency, combined-failure behavior, canary volume and rollback target. A passing model-only metric does not cover a changed extractor. Approve only the tested chain combination, and make rollback restore the whole chain pointer. If an independent stage must be promoted, rerun the end-to-end fixture set and record a new chain revision. Preserve any unreadable receipts for manual handling rather than replacing them with an apparently clean score.
Implementation
def chain_release_gate(stages, expected, fallback_load, capacity):
mismatches = [name for name, digest in expected.items()
if stages.get(name) != digest]
if mismatches:
return {"state": "hold", "reason": "stage-identity",
"stages": mismatches}
if fallback_load > capacity:
return {"state": "hold", "reason": "review-capacity"}
return {"state": "canary"}
expected = {"extractor": "extract-47", "risk": "risk-82",
"policy": "policy-3"}
assert chain_release_gate(expected, expected, 18, 23)["state"] == "canary"
assert chain_release_gate({**expected, "extractor": "extract-46"},
expected, 18, 23)["reason"] == "stage-identity"
Performance and operating cost
Comparing s stage identities takes O(s) time and O(s) mismatch space. End-to-end tests have cost proportional to the number of fixtures and stage invocations; a live canary adds staff and fallback capacity. The gate cannot substitute for stage outputs, deadline measurements or evidence that the fallback actually reaches the review queue.
Common Mistakes
- Approving a model digest without pinning its extractor.
- Starting later stages after an earlier required stage fails.
- Letting a degraded route overload human review.
- Rolling back only the risk model and retaining the new extractor.
Read next
- Multi-stage inference: pin each stage and its contract
- Test partial failure and deadline exhaustion in model chains
- Project: operate receipt scoring with a deadline and overload path
- Project: release a versioned receipt feature admission gate
- Inference logs: keep diagnostic joins without copying sensitive payloads
