Writer fencing rejects commits from a former primary after another region takes ownership; source positions let the replacement resume without a gap or double publication.
Failover writer fencing and source positions
Give each promotion a higher epoch
A coordinator issues monotonically increasing writer epochs from durable control storage. Every table commit carries its epoch, and the commit gate rejects an older one. A lease timeout alone is not enough: a paused old worker can wake after its lease expires and still send writes. The epoch makes that stale writer observable and rejectable.
Promote from a complete generation
The recovery region first selects one manifest whose data files, schema IDs, state and source positions are present. It must not resume from the newest source offset if the table snapshot reflects an older offset. Replaying the intervening events under stable event IDs should reconstruct the missing result. Envelope identity supplies the deduplication boundary.
Separate duplicate delivery from duplicate effects
After promotion, some source records may be delivered again. This is acceptable if the sink’s business key and version rule reject repeated effects. A side channel such as an email or billing call needs its own idempotency contract. Sink replay cannot be inferred from the table commit alone.
Handle the returning primary
When the old region recovers, it must not automatically resume as writer. Reconcile its last epoch and committed snapshot with the promoted region, then either discard candidate outputs or replay missing source intervals into the new leader. A second promotion back is a controlled release with a still higher epoch. Keep both branches for audit until divergence is resolved.
Inject delayed acknowledgements
Pause the primary immediately before commit, promote the secondary, commit one new generation, then release the old request. The old commit must fail because its epoch is lower. Validate one current table pointer, no missing source positions and unchanged financial totals. Record the rejected attempt and candidate-file cleanup path.
Implementation
published = {"epoch": 8, "source_position": 470, "total_cents": 8200}
def fenced_commit(current, epoch, source_position, total_cents):
if epoch < current["epoch"]:
raise PermissionError("stale writer epoch")
if source_position < current["source_position"]:
raise ValueError("source position moved backward")
return {"epoch": epoch, "source_position": source_position,
"total_cents": total_cents}
promoted = fenced_commit(published, 9, 471, 9100)
try:
fenced_commit(promoted, 8, 472, 10200)
except PermissionError:
pass
else:
raise AssertionError("former primary committed")
assert promoted["source_position"] == 471Performance and operating cost
The reference gate is O(1) time and space. A real implementation must make epoch comparison and table pointer update atomic in durable storage; a read followed by an unguarded write is unsafe. Replicated source logs, candidate files and promotion drills add storage and compute, but prevent two regions from publishing competing histories.
Common Mistakes
- Do not trust a time-based lease without a stale-writer commit check.
- Do not resume from an offset newer than the recoverable table generation.
- Do not let a returning primary reclaim writer status automatically.
