Choose the next recovery action from evidence of user impact, writer authority, replica position, and client behavior. The goal is a functioning service with one accepted data history and a tested path through the next failure.
Review these lessons
- Recovery-time budgets: measure every stage of service return
- Standby admission: prove the recovery region can accept real work
- Disaster declaration: define authority, scope, and stop conditions
- Failover ordering: fence, promote, validate, then move clients
- Write-gap ledgers: account for acknowledged work that never reached the standby
- Cutover drains: handle old connections and ambiguous client retries
- Failback: rebuild the former primary before returning write authority
- Recovery drills: retain evidence and reset the next line of defense
Other checks
Common Mistakes
- Do not confuse traffic movement with write fencing.
- Do not infer a safe failback from a restarted old database.
