Simulate a customer-visible scoring failure, identify the responsible release and prove the chosen mitigation helps.
Project: run a receipt-model incident drill with honest mitigation
Prepare the drill
The receipt service has a stable model digest, a recent candidate and a feature producer that emits amount in the wrong unit for one region. Introduce a sustained timeout burst as a separate fault. Freeze a timeline with deployment and symptom timestamps, plus request counts by region and outcome. Alert rules should page on customer impact while keeping a drift-only signal in the investigation queue.
Require a bounded diagnosis
The responder identifies which region and requests are affected, checks feature admission failures, compares model digests and verifies the old artifact remains loadable. A model rollback is a candidate action, not a default answer. The wrong-unit fault needs a producer repair or rejected-input fallback. The timeout burst needs capacity or dependency mitigation. The evidence timeline forces each claim to carry an observation window and release identity.
Execute and verify mitigation
Route malformed features to approved manual review, stop the faulty producer release and verify that the feature contract accepts new traffic. Bound the inference queue for the timeout incident and return a declared fallback before the deadline. If the candidate model separately fails a quality or safety guardrail, restore the old model and compatible producer together. Record actions and compare pre/post customer outcomes on the same population.
Write the review artifact
Report time to detection, time to customer mitigation, the first incorrect hypothesis, root cause, recovery tests and follow-up owner. Keep metric dimensions low-cardinality and sensitive receipt payloads restricted. Verify the next on-call person can follow the runbook without access to the original responders. The project is complete when the drill catches an ineffective rollback and demonstrates a mitigation that actually changes the customer symptom.
Implementation
def mitigation_result(before, after):
if before["requests"] < 470 or after["requests"] < 470:
return {"state": "inconclusive", "reason": "small-window"}
old_rate = before["failed"] / before["requests"]
new_rate = after["failed"] / after["requests"]
if new_rate >= old_rate:
return {"state": "ineffective", "before": old_rate, "after": new_rate}
return {"state": "improved", "before": old_rate, "after": new_rate}
before = {"requests": 500, "failed": 39}
after = {"requests": 500, "failed": 7}
assert mitigation_result(before, after)["state"] == "improved"
assert mitigation_result(before, before)["state"] == "ineffective"
Performance and operating cost
The compact comparison is O(1) time and space. Incident analysis costs retained metrics, trace sampling and operator time. A lower post-action failure rate alone does not prove the action caused improvement when traffic mix changes; compare slices and timelines before attributing the result.
Common Mistakes
- Choosing model rollback before inspecting the feature contract.
- Using an unapproved safe score as an outage fallback.
- Comparing pre/post windows with different traffic populations.
- Ending the drill without an owner for the root-cause repair.
Read next
- Model alerts: page on customer symptoms with a named owner
- Model incidents: build a release and evidence timeline before rollback
- Feature contracts: admit only usable inference records
- Serving overload: bound queues and choose a fallback before time runs out
- Shadow and canary rollout: compare a candidate without losing a rollback
Continue the workflow: Project: fail over receipt inference and reconcile every decision.
Continue the workflow: Project: measure receipt inference without losing the denominator.
Continue the workflow: Project: stage a receipt-quality model on offline handhelds.
