The evidence package links a declared scenario to the actual user-path result, timestamps, data gap, operator decisions, artifacts, and stop-condition behavior. It distinguishes a tabletop, a synthetic fault in an isolated environment, and a production rotation; passing one does not prove the others. Reset includes restoring standby replication, backups, routing controls, credentials, alerts, capacity, and queue ownership. A drill that leaves the standby stale has improved confidence on paper while reducing real recovery readiness.
Recovery drills: retain evidence and reset the next line of defense
Operational decision
A catalog-search team runs a scoped region-rotation drill for a disposable tenant. The recorder captures start and end markers from an external probe, one successful catalog update, a rejected stale-region update, replica replay positions, DNS and connection observations, capacity peak, and every manual override. A stop threshold fires when synthetic checkout errors exceed the agreed bound; the team halts fault injection, checks residual effects, and records whether the stop occurred before or after customer-path impact. It then reverses the exercise with a separately reviewed failback sequence, re-establishes replication, verifies a recent restorable backup and key access, restores old routing-control defaults, and checks alerts in both regions. The report labels untested paths, such as a full provider-region loss, instead of extending a narrow drill result to them. Findings become owned actions with deadlines and a repeat-test condition. The next drill uses a changed assumption, such as a clean node with no cached image, to reveal a different dependency rather than replaying the same happy path indefinitely.
Catalog recovery evidence
Scenario and blast radius recorded
External user-path start/end timestamps retained
Old-writer rejection and new-writer success observed
Replay point and actual data gap captured
Stop threshold and residual effects inspected
Failback and standby protection restored
Open findings have owner and repeat-test conditionCost and verification
Evidence collection is O(E) for E events and artifacts, while running two regions and restoring copies drives the real cost. Retention should be long enough to compare drills and incident outcomes without keeping sensitive payloads unnecessarily. Measure how often the drill meets its declared recovery targets, how long protection remains degraded afterward, repeat findings, and whether action owners actually close their gaps before the next exercise.
Common Mistakes
- Do not call a tabletop exercise proof of live failover timing.
- Do not stop recording at first successful request while protection remains degraded.
- Do not leave routing controls or replication in their temporary drill state.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Incident reviews: turn a timeline into tested corrective work
- Release evidence: tie one deployed digest to one approval decision
- Backups and disaster recovery: prove the restore path
- Failback: rebuild the former primary before returning write authority
