An incident response begins by establishing scope and stopping harm. Build a timeline from user-visible failures, release changes, service events, queue lag, and verified recovery checks. Timestamps from different systems can be skewed; sequence and operation IDs often give stronger ordering evidence than wall time alone. Containment may mean disabling a feature path, pausing a worker, reducing traffic, or rolling back a release. A rollback is safe only when its old code can still read current data and when queued jobs have a compatible consumer. Keep the decision, expected effect, and reversal condition visible to responders.
Incident Containment and Evidence Timeline
Working case
Release 29 changes the case receipt worker. At 14:07, completed-save failures rise; at 14:09, queue lag grows; by 14:12, reviewers have 62 unresolved receipts. The API still accepts new saves. Disabling only the receipt feature prevents new incomplete promises while preserving read access. The team checks whether release 28 can process messages written by release 29 before rolling back. A timeline lists the first bad operation ID, deployment event, containment action, backlog size, and the synthetic case-save result that proves recovery. A quiet error graph alone is not proof that pending work cleared.
Implementation boundary
function containmentChoice(currentWritesSafe, previousReaderCompatible) {
if (!currentWritesSafe) return "pause-writes";
return previousReaderCompatible ? "rollback-candidate" : "isolate-worker";
}
console.log(containmentChoice(true, false));
// Output: isolate-workerThe helper is a decision sketch, not automatic incident control. A responder should verify user impact and blast radius, identify the suspect change, choose the smallest reversible action that stops further damage, and assign an owner to pending operations. Record evidence before changing systems when doing so will not delay containment. Test rollback compatibility against schema and queued-message versions; an old worker may misread new payloads. After containment, run the exact user journey that failed, confirm terminal operation outcomes, inspect backlog drain, and monitor for recurrence. A post-incident record should separate observed facts from hypotheses.
Cost and boundaries
A timeline adds O(n) review work for n relevant events; filtering to bounded operation and release identifiers keeps it tractable. Pausing a worker preserves data but increases queue age and storage; dropping messages reduces backlog at the cost of lost work and must have an explicit recovery rule. A rollback can restore behavior quickly yet fail if migrations or contracts are not backward compatible. A feature switch can contain one path but adds operational state that must be audited. Practice these actions ahead of an outage and measure time to detect, contain, and verify—not only time until a dashboard turns green.
Failure trace
The team sees a spike after release 29 and immediately deploys release 28. The older worker reads a new message field incorrectly and marks 62 receipts complete without writing them. The apparent rollback deepens the incident. Check message and schema compatibility first, then isolate the worker if necessary. Another failure closes the incident when errors stop, ignoring accepted jobs already stuck in the queue. Include backlog and terminal operation counts in recovery evidence. If an action has no clear reversal condition, it can remain enabled for weeks and create a second failure mode.
Verification
- The timeline distinguishes facts from tentative causes.
- Rollback is gated by schema and queued-message compatibility.
- Recovery evidence includes pending operation outcomes, not just fresh traffic.
Practice drill
Inject a worker decode error after 29 accepted case saves. Write a timeline with first impact, release, queue lag, containment, and confirmed recovery. Attempt a rollback using an incompatible message fixture and verify the runbook blocks it. Pause writes only for the affected operation and keep unrelated reads available. Drain the backlog into a safe consumer, then check every accepted operation reaches a terminal state or explicit repair queue. Repeat with clock skew between API and worker logs, using sequence and operation IDs to reconstruct order.
Decision note
Contain the smallest failing path, then verify both new journeys and the backlog created before containment.
Common Mistakes
- Rolling back code without checking current data compatibility.
- Treating a quiet error chart as proof that backlog is repaired.
- Leaving a containment switch without an owner or reversal condition.
Connected lessons
Production Signals and Incident Decisions; Trace Context Across Requests and Jobs; User Journey SLOs and Burn Alerts; Telemetry Shapes, Redaction, and Cardinality; Release checks: prove the critical route and prepare a rollback; Expand-and-Contract Schema Migrations; Accepted Operations and Status Resources.
Apply and check
Build Project: case-save incident evidence and review Web Development: operations and sync decisions quiz.
