An asynchronous database replica applies committed changes after the primary accepts them. The delay between those events is replication lag. A query routed to the replica can therefore return an older value even though the write succeeded. Lag is a read-consistency issue during normal operation and a possible data-loss window when a lagging replica is promoted after primary failure.
Replica lag: define when a read is allowed to be stale
Operational decision
A claims API writes a settlement state and immediately displays it to the operator. Route the confirmation read to the primary or require a read position at least as recent as the write; a generic read pool can show the previous state. The SQL fragment reads replication position on a PostgreSQL primary and replay position on a standby, so run each statement against its named role rather than as one transaction. Monitor both byte distance and elapsed replay age, then inject delayed replication in a disposable system and observe the user path. Before promoting a standby, record the last confirmed write and the replay position on the candidate. After promotion, reconcile external acknowledgements with the new primary; an asynchronous replica may lack acknowledged transactions. State a recovery-point objective in terms of business records, not only seconds of lag. A healthy connection to the standby does not establish that it has caught up.
-- Run on the primary
SELECT pg_current_wal_lsn() AS primary_lsn;
-- Run separately on the standby
SELECT pg_last_wal_replay_lsn() AS standby_replay_lsn;Cost and verification
Reading from replicas can reduce primary query pressure but adds lag measurement, routing rules, and failover reconciliation. Stronger synchronous durability can increase commit latency or block writes when a required standby is unavailable; the exact trade depends on configuration. Position alone does not prove a business effect was applied correctly, so retain an end-to-end confirmation check. Alert on sustained lag relative to the service's tolerance and on a stalled replay process.
Common Mistakes
- Do not route immediate read-after-write confirmation to an unconstrained replica.
- Do not equate a reachable standby with a caught-up standby.
- Do not promote without recording the potential acknowledged-write gap.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Database pool pressure: bound waiting before the database collapses
- Multi-region failover: define write ownership before moving traffic
- Backups and disaster recovery: prove the restore path
- Synthetic transactions: measure the route a user actually takes
