Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Regional Failover, Fencing, and Replay

Last updated: 5 Oct 20267 min read
tutorial
IntermediateBy AITrove Editorial

A failover transfers write authority. If the old region can still accept writes while the new region starts writing, records may diverge even if traffic has mostly shifted. Fencing is the mechanism that prevents an older leader from continuing work after a newer generation takes authority. Asynchronous replication also means acknowledged writes may not yet exist in the recovery copy. A recovery plan must state a recovery-point budget, a recovery-time budget, the authority-generation rule, and how uncertain requests and queued jobs are replayed.

Working case

Region 29 owns case writes and region 47 holds a recovery replica. Region 29 stops responding to health checks after case 62 receives revision 47. The operator checks replica position; region 47 has only revision 46. The team pauses new writes, fences the old writer, records the possible missing revision, then promotes region 47 under a new generation. A browser retry carries the same idempotency key so the case change does not apply twice. Queued email and export work are inspected against durable outbox records rather than restarted blindly.

Implementation boundary

javascript
function acceptsWrite(writerGeneration, activeGeneration) {
  return writerGeneration === activeGeneration;
}
console.log(acceptsWrite(6, 7));
// Output: false

Use the chosen data service’s documented promotion and fencing controls; a DNS change alone is not fencing. Persist an authority generation or lease where the architecture requires it and make writers reject stale generations. During promotion, compare replication positions and record the possible loss window. Make mutating requests idempotent, especially when the client cannot tell whether a timed-out write committed. Resume jobs from durable intents and reconcile side effects such as email before retry. Practice failback separately; restoring the old region without resynchronizing it can reintroduce divergent data.

Cost and boundaries

A passive region costs standby storage and periodic drills; an active region adds steady compute, network replication, and more difficult write coordination. Fencing and promotion may reduce availability for a short period, but unchecked dual writers can cause permanent data conflicts. Replication lag sets a possible recovery-point loss for asynchronous designs. Measure time to detect, time to fence, lag at promotion, idempotency reuse, replay count, and reconciliation outcomes. Do not promise zero loss when the chosen storage cannot provide it.

Failure trace

Traffic moves to region 47 but a batch worker in region 29 keeps writing case assignments, producing two revision 47 records. Another retry invents a new request key and sends duplicate notification email. Test old-region partial availability, stale worker credentials, a partition that leaves both regions alive, replica lag, a write acknowledged just before failure, duplicate outbox delivery, and failback. The drill should prove which authority generation accepted each write and which side effect still needs human reconciliation.

Verification

  • Old-generation writers reject work after promotion.
  • Possible unreplicated writes are accounted for.
  • Retries and jobs reuse durable identities.

Practice drill

Start writer generation 6 in region 29, then simulate a partition and promote generation 7 in region 47. Submit one old-generation write and assert rejection. Replay a timed-out case save with its original idempotency key. Compare acknowledged revision 47 with a recovery replica at 46 and record the resulting recovery-point gap. Reconcile a queued email intent before enabling outbound delivery in the new region. Repeat failback only after data is resynchronized.

Decision note

Transfer write authority with fencing and explicit loss accounting before resuming mutations.

Common Mistakes

  • Using traffic routing as the only writer fence.
  • Promising zero data loss with asynchronous replication.
  • Replaying external side effects without reconciliation.

Connected lessons

Build Project: two-region case service recovery and review Web Development: motion and regional state decisions quiz; follow Multi-Region Web State and Recovery; Regional Routing, Static and Private Cache Boundaries; Replica Lag, Read-Your-Write, and Version Cursors; Regional Data Location and Operational Evidence; Idempotent Write Requests and Lost Responses; Backup and Restore Drills.

web-tech
web-development
Storage details