Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: rehearse a fenced regional recovery and failback

Last updated: 5 Oct 202610 min read
project
AdvancedBy AITrove Editorial

Run this drill in two disposable environments with a synthetic reservation service, a database replica, a queue, and a client probe outside both regions. Use a small preloaded dataset and a stable request identifier for each reservation. The exercise passes only when a customer-path reservation works in the recovery region, the former writer rejects new commits, uncertain acknowledgements are reconciled, and a protected standby exists after failback.

Declare the scenario and budget

Set a workload RTO and RPO before starting. Define the recovery clock's start and end events, maximum synthetic error rate, stop operator, allowed tenant, and hard block for an unconfirmed write fence. Capture external-probe timestamps rather than relying on clocks inside the impaired region. Prove the standby can pull the tested image on a clean node, decrypt required secrets, connect to its dependencies, and scale within quota. Fail one admission gate deliberately and show that the operator refuses promotion until it is repaired or an approved degraded mode is recorded.

Inject failure and move authority

Write several reservations near the fault boundary; arrange for one acknowledged request to be absent from the standby and one client response to be lost after commit. Keep their identifiers and external side effects observable. Fence the primary using a monotonic epoch or an equivalent enforced mechanism, then prove a stale worker cannot write. Record the standby replay point. Promote one standby, run a synthetic write/read through the application, and shift new traffic. Keep old connections alive long enough to test stale-write refusal and same-key retries. If the promotion command times out, inspect actual state before trying again.

Output
Recovery drill acceptance
RTO start and end are external customer-path events
Standby admission includes clean-node pull and capacity
Old writer rejects an outdated epoch before traffic moves
Promoted writer accepts one isolated synthetic reservation
Uncertain requests have acknowledgement and side-effect status
Old connections cannot create a second committed reservation
Former primary follows the promoted history before return
Backup, replica, alerts, and routing controls reset after drill

Reconcile and fail back

Build a per-request write-gap ledger from client acknowledgements, database positions, queue deliveries, and any external provider IDs. Do not replay an ambiguous operation until its original side effect is known. Bring the former primary back only behind a write fence, preserve a copy for investigation, and either resynchronize it to the promoted timeline or rebuild it from a clean base. Test replication catch-up and a second application-path cutover. Restore a standby in the opposite direction, verify recent backup readability and key access, and check that alerting and routing controls have returned to their intended state.

Report the result

Report elapsed time by stage, actual data gap, customer errors during convergence, old-region rejected writes, peak standby saturation, stop-condition behavior, manual overrides, and time spent without a healthy standby. Label environment-specific results as such: a local simulation does not establish provider-region behavior or production throughput. Assign an owner and repeat-test condition to every failed gate. A repeat drill should remove one hidden dependency, such as cached image layers or a preexisting connection.

Common Mistakes

  • Do not switch traffic while two writers can commit.
  • Do not replay a transaction with unknown external side effects.
  • Do not leave the former primary writable or the new region without a standby.

Connected lessons

devops
project
Storage details