Prepare an incident-communication packet for fictional Harbor Checkout. The checkout API snapshot HC-1410 covers 14:00-14:10 UTC: 3,720 failures among 18,600 attempts, exactly 20% of measured attempts. The project produces a public draft, an internal handoff, a recovery gate, and a correction route. It does not publish a status message or operate the checkout service. An incident commander must approve the exact external wording after the metrics lead verifies its source.
Project: draft Harbor Checkout incident updates
Draft the first update
State that elevated checkout failures are under investigation. Name the affected API path and measured window. Do not translate attempts into unique customers, claim that every transaction failed, assign a root cause, or offer a restoration time without evidence. Put the next update at 14:30 UTC. A recent deployment is an internal hypothesis, not a public fact. The communications lead drafts from HC-1410; the incident commander checks the text and a human publisher controls the status page.
Record action without declaring recovery
At 14:16 UTC an operator starts a rollback. The handoff records its receipt and the next owner, but it keeps customer recovery open. A later window must return to the agreed baseline, and a synthetic checkout must succeed before the commander approves a resolution draft. If the metric definition changes after publication, timestamp a correction with the revised numerator, denominator, and reason. The incident review retains the earlier draft, approvals, receipts, corrected impact, and owners for follow-up work.
HC-1410: 3,720 failed / 18,600 checkout API attempts = 20%.
14:12 UTC public draft: investigating measured checkout failures.
14:16 UTC rollback starts; recovery is not yet confirmed.
14:20 UTC handoff: receipt, owner, open checks, 14:30 update.
Resolution: later metric window plus synthetic checkout plus approval.
Correction: timestamp any changed count; preserve earlier record.Performance and operating cost
Counting N raw events is O(N) time and O(1) aggregate memory; validating counter definitions, ingestion, and retries costs additional operator work. A handoff over A active actions is O(A) review. The two recovery checks test different failure modes, while the observation window catches a relapse that a single synthetic request could miss. Approval delays publication slightly, but prevents an unverified draft from becoming a customer-facing claim. Use compact evidence IDs instead of copying sensitive logs into every prompt.
Common Mistakes
- Do not call the 20% attempt rate a customer rate.
- Do not publish the internal deployment hypothesis as root cause.
- Do not mark the service restored when the rollback only starts.
- Do not erase a wrong update; issue a timestamped correction.
Connected lessons
- Incident updates: audience, owner, and release gate
- Incident impact: window, denominator, and scope
- Incident updates: separate facts, hypotheses, and ETA
- Incident handoffs: preserve actions and decisions
- Incident resolution: verify recovery and correct the record
- Incident triage prompts: build a timestamped evidence ledger
- Aggregation prompts: pin the denominator and recompute the rate
- Incident response: contain impact, then learn
- Incident reviews: turn a timeline into tested corrective work
- Incident-communication prompt decisions
