Recovery time for a deployment-caused failure needs an explicit impairment start and a verified restoration endpoint. Deployment activation, first user impact, alert detection, operator acknowledgement, rollback command, and successful public probe are separate timestamps. Choose and publish the clock used in the report; retain the other timestamps to explain detection and mitigation delay. Exclude unrelated infrastructure incidents from this change-specific measure while still tracking them in the broader incident program. If impact start is uncertain, record a bounded interval and label the measurement uncertain. A code rollback can finish before data repair or cache expiry restores the actual journey.
Failed deployment recovery time: keep impact, detection, and restoration clocks distinct
Operational decision
A faulty release activates at 14:02. The first failed checkout occurs at 14:06, the alert fires at 14:14, the team starts rollback at 14:21, and a synthetic checkout plus real request sample pass at 14:55. Under a first-impact-to-verified-restoration rule, recovery takes 49 minutes. Detection lag is 8 minutes and the post-detection recovery interval is 41 minutes. A dashboard that turns green at 14:34 is not the endpoint if payment confirmation remains broken. Preserve the release ID, affected cohort, probe evidence, and timestamp sources so the incident review can reconstruct the interval.
Activation: 14:02
First verified user impact: 14:06
Detection: 14:14
Rollback started: 14:21
User path restored: 14:55
Failed-deployment recovery time: 49 min
Detection lag: 8 min; post-detection repair: 41 minCost and verification
Computing a duration per incident is O(1) after its evidence is assembled, but collecting trustworthy start and end times can require traces, synthetic checks, and customer reports. Report the distribution across incidents, including long tails and uncertain records. Avoid an average when a few severe events dominate user harm. A fast rollback should not appear as fast recovery if the old version cannot read candidate-written data or a stale edge cache still serves the bad page.
Common Mistakes
- Do not start every recovery clock at alert detection without saying so.
- Do not stop the clock when an operator presses rollback.
- Do not mix deployment-caused failures with unrelated outages in this measure.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Change fail rate: attribute immediate intervention to a production deployment
- Synthetic transactions: measure the route a user actually takes
- Serverless rollback: restore code traffic without assuming data rolled back
