Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Write-gap ledgers: account for acknowledged work that never reached the standby

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

The recovery-point objective is a business loss bound, not simply a replication-lag metric. Compare the last durable primary acknowledgement with the selected standby replay point and account for writes in API logs, idempotency records, transactional outboxes, queues, external payment systems, and object stores. Those systems may have different replication boundaries. Missing and duplicate work require different actions: reconcile absent accepted requests, while deduplicating retries that arrive after the client lost its response.

Operational decision

A payout service fails after acknowledging transfers numbered 641 through 648. The standby contains records through 645. The operator freezes new disbursements and creates a ledger with each request ID, client acknowledgement, database commit position, outbox event, queue delivery, and bank-provider reference. Requests 646 and 647 reached the provider but not the standby, so replaying them as new transfers would double-pay; reconstruct them using provider IDs and the original idempotency keys. Request 648 has a client acknowledgement but no provider receipt, so it needs a bounded investigation and customer-facing state, not a guessed success. A queue in the lost region might retain undelivered messages when service returns; keep its consumers paused until ownership and deduplication rules are set. Record the earliest unrecoverable acknowledgement and compare its time with the declared data-loss budget. A zero-byte replica lag reading after promotion says nothing about writes that never arrived. Have finance or the domain owner sign off the disposition of each uncertain transfer before normal automation resumes.

Output
Payout write-gap record
Request: 646
Client: acknowledged
Standby: absent at selected replay point
Outbox: present only in old region
Provider: accepted with stable reference
Action: reconstruct local state; never resubmit payment
Owner: payout operations
Status: pending independent reconciliation

Cost and verification

The ledger needs O(W) review for W writes in the uncertainty window; indexed request IDs and provider references reduce search cost but do not remove domain review. Retain raw evidence long enough to resolve disputes, while restricting sensitive financial fields. Measure unresolved acknowledgements, duplicate-provider submissions, outbox gaps, reconciliation age, and the actual time between last accepted primary write and standby recovery point.

Common Mistakes

  • Do not equate low replica lag with zero lost acknowledged writes.
  • Do not replay a transfer until external side effects are checked.
  • Do not restart old-region consumers before queue ownership is settled.

Connected lessons

Practice and check

devops
disaster-recovery
Storage details