A recovery-time objective bounds the elapsed time between a defined incident start and a defined service-restored condition. Split that interval into detection, declaration, fencing, data promotion, capacity, traffic movement, and user-path verification. The stages may overlap, but record actual timestamps and dependencies rather than adding guessed durations. A database responding to a local probe is only one checkpoint; an API that still cannot authenticate or settle a payment has not met its workload RTO.
Recovery-time budgets: measure every stage of service return
Operational decision
A claims-submission service declares a regional outage at 09:14. Its recovery record captures the last known successful customer submission, alert timestamp, human decision, primary-write fence confirmation, standby replay point, promotion, application readiness, first successful synthetic claim, and restoration of full intake. The team chooses the first successful end-to-end claim as the recovery endpoint and records degraded throughput separately. During a rehearsal, authentication depends on a regional key service that was omitted from the original timeline; the team adds that dependency and its alternate-region readiness check. Measure the critical path, not just the sum of every task, because capacity scaling can proceed while the write fence is established. Do not reset the stopwatch when a new incident commander takes over. If the recovery objective is missed, retain the measured interval and the stage that consumed the budget. Repeated drills reveal whether detection, operator decision, provisioning, replica catch-up, or traffic convergence is the binding delay. State the acceptance journey before the drill so the result cannot be redefined after a slow recovery.
Claims recovery clock
Start: last confirmed healthy customer claim at 09:14
Declare: incident decision recorded
Fence: former primary rejects writes
Data: selected standby replay point recorded
Ready: application and dependencies pass checks
Traffic: customer endpoint routes to standby
End: new claim accepted and queryable
Separate: time to full planned throughputCost and verification
Capturing T stage events is O(T) in record size and small beside the recovery itself. Parallel work shortens wall time only when it does not violate ordering: traffic must not reach an unfenced second writer. Measure p50 and worst rehearsed recovery times, stage variance, stale timestamps, and customer-path failures after a nominally successful promotion. Reserve time for verification and a retry instead of allocating every minute to provisioning.
Common Mistakes
- Do not stop the clock when only the database is writable.
- Do not sum overlapping stage durations as if they were sequential.
- Do not change the accepted recovery endpoint after seeing the drill result.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Backups and disaster recovery: prove the restore path
- Multi-region failover: define write ownership before moving traffic
- Incident response: contain impact, then learn
- Release evidence: tie one deployed digest to one approval decision
