A PostgreSQL replication slot retains log records that a replica or logical consumer may still need. When a consumer stops advancing, the oldest required WAL position can stay fixed while writes continue, increasing disk use. A retention ceiling can protect the primary's disk, but if required WAL is discarded the consumer may be unable to resume from its prior position. Dropping a slot relieves retention at the cost of that replay path.
Replication slots: bound retained WAL before a consumer outage fills the disk
Operational decision
A receipt change-feed connector is offline during a promotion weekend. Query slot activity, restart_lsn, confirmed_flush_lsn, wal_status, and bytes between the current WAL position and each slot's restart position. The SQL is read-only and targets a PostgreSQL version that exposes those columns. Compare retained bytes with free storage, current WAL generation rate, and the expected repair time; an absolute byte count without a growth rate gives no time-to-full estimate. Restore the connector if its offset and slot still align, then verify event counts and a receipt marker across the catch-up window. If the slot is lost or the consumer offset no longer exists in WAL, stop the connector and plan a controlled resnapshot with sink deduplication before deleting the old slot. Never drop an apparently inactive slot just because its client process is absent; it may be the only path to unconsumed changes. Record the ownership and recovery decision for every slot.
SELECT slot_name, slot_type, active, restart_lsn, confirmed_flush_lsn, wal_status,
pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn) AS retained_bytes
FROM pg_replication_slots
ORDER BY retained_bytes DESC NULLS LAST;Cost and verification
WAL retention buys recovery time but spends disk. A strict cap protects the primary while increasing the chance of resnapshot and downstream reconciliation. Queries over slot metadata are cheap relative to a full table scan, but regular alerting should include disk headroom and write rate. A consumer appearing connected does not prove its confirmed position is moving. Recheck downstream business effects after catch-up, not only slot activity.
Common Mistakes
- Do not drop an inactive slot before identifying its consumer and offset.
- Do not use free disk bytes alone without WAL growth rate and repair time.
- Do not call a reconnected consumer recovered until its sink catches up.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Replica lag: define when a read is allowed to be stale
- Backups and disaster recovery: prove the restore path
- Node pressure eviction: trace lost Pods to exhausted local resources
- Observability: join metrics, logs, and traces
