Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Failover fencing: prevent two writable database leaders

Last updated: 7 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

Failover fencing removes write authority from a failed or isolated leader before a replacement accepts writes. It matters when a network partition makes a healthy server appear dead to one observer while clients can still reach it. A promotion that changes only the connection string can leave two writable databases, each accepting valid-looking transactions that cannot be reconciled by replay alone.

Operational decision

A settlement database has a primary in one zone and a standby in another. Before promotion, the operator records the last acknowledged payout ID and the candidate's replay position. Stop the former primary's write endpoint with a tested mechanism outside the application, such as a managed-service promotion protocol or host/network isolation that the old process cannot bypass. The text gate is an operational record, not an instruction to perform blind node shutdown. Confirm the old endpoint rejects a test write, then promote and route one synthetic payout through the new writer. If the former host returns, keep it isolated until its data is compared and it is rebuilt as a replica; do not reconnect it as a peer with divergent writes. In an asynchronous system, separately reconcile transactions acknowledged on the old primary but not replayed on the promoted candidate. A consensus store may reject minority writes, but that property must be verified for the actual database and topology rather than assumed from the presence of a lease.

Output
Settlement failover gate
Old writer: isolated from application and direct clients
Candidate: replay position and missing-write window recorded
Negative test: old endpoint rejects a controlled write
Promotion: one named operator and one candidate
Positive test: new writer commits a synthetic payout
Return: old host rebuilt as replica after divergence check

Cost and verification

Fencing can lengthen recovery because the team must establish which side has authority before traffic moves. That cost prevents much larger inconsistency from simultaneous writers. An aggressive automatic failover can improve apparent availability while losing acknowledged writes if the candidate was behind. Measure both recovery time and the gap between the last acknowledged business operation and the new writer's state. Test direct clients, scheduled jobs, and bypass routes; changing a load balancer alone may leave an old write path open.

Common Mistakes

  • Do not promote based only on a missed heartbeat while the old writer may accept traffic.
  • Do not rejoin a divergent former primary without rebuilding or reconciling it.
  • Do not call zero split-brain risk proof that no acknowledged data was lost.

Connected lessons

Practice and check

Advanced follow-up

Recovery execution follow-up

Redis operating follow-up

devops
operations
Storage details