Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Redis Sentinel failover: reconcile acknowledged writes across a new primary

Last updated: 2 Oct 20267 min read
tutorial
AdvancedBy AITrove Editorial

Redis Sentinel monitors primary and replica instances and helps clients discover the current primary after failover. Redis replication is asynchronous, so a primary can acknowledge a write before a replica has it. If a lagging replica is promoted, a recently acknowledged write can disappear from the new primary. The old primary may later rejoin as a replica, replacing its divergent state. Sentinel is therefore an availability mechanism with an explicit write-gap risk, not a substitute for an authoritative transaction log. A client must discover the new primary through Sentinel-aware behavior and handle retries without inventing new business identities.

Operational decision

A reservation service stores a temporary availability view in Redis but commits the actual reservation in a database. At 09:47, a primary failure occurs after Redis acknowledges token 7284; the promoted replica lacks it. The service reconstructs the view from the database ledger and rejects any action that relies only on the missing cache key. A disposable drill pauses replica traffic, issues numbered writes, kills the primary, and compares acknowledged IDs with the promoted dataset. It also reconnects clients through Sentinel discovery and confirms that stale connections do not continue to drive authoritative work on an old endpoint. Measure time to promotion, missing-write count, and database reconciliation time separately.

Output
Primary acknowledged: token-7284
Replica before fault: through token-7282
After promotion: reconcile token-7283 and token-7284
Authority: reservation database ledger
Client action: discover current primary and retry by stable request ID
Recovery evidence: missing IDs and rebuild duration

Cost and verification

Replication lag and failover time depend on load and network conditions. More replicas improve available promotion choices but do not make asynchronous writes atomic. A command such as WAIT can reduce the chance that an acknowledged write is absent from replicas, yet it does not create a strong-consistency guarantee or protect against every persistence and failover sequence. Monitor replica offsets, Sentinel quorum, client reconnect errors, and old-primary fencing. If losing one acknowledged write is unacceptable, choose an authority with the required durability model instead of describing Sentinel as lossless.

Common Mistakes

  • Do not promise lossless failover from asynchronous replicas.
  • Do not treat a new primary address as proof that the application state reconciled.
  • Do not retry an ambiguous reservation with a fresh business ID.

Connected lessons

Practice and check

devops
redis
data-operations
Storage details