Run this exercise in a disposable PostgreSQL environment with synthetic payouts, a change-feed connector, an outbox dispatcher, and a search sink keyed by stable business IDs. Take a restorable base backup before injecting faults. Capture initial WAL generation rate, replication-slot positions, sink checkpoint, oldest outbox age, active transaction ages, and disk headroom. Keep the experiment separate from a shared or production database.
Project: recover a database change stream and restore boundary
Interrupt the change stream
Pause the connector while payouts continue, then calculate retained WAL growth and the time remaining before the disk budget is exhausted. Resume it before its slot is lost and reconcile source and sink IDs. Rebuild a second sink with a snapshot while one payout updates and another is deleted; crash the connector just before it records snapshot completion. After restart, the sink must have neither missing nor resurrected rows. Settle one payout and insert its outbox event in a single transaction. Crash the dispatcher after publishing but before checkpointing; the consumer must produce one durable business effect for that event ID.
Data continuity acceptance gates
Slot: restart position advances; retained WAL returns to budget
CDC: source and sink IDs and sampled values reconcile
Outbox: one transition, one event identity, duplicate delivery safe
Index: valid catalog state and measured query improvement
DDL: blocked migration fails inside lock budget
Vacuum: old transaction owner found and horizon advances
Archive: detached partition restores and queries correctly
PITR: target boundary and acknowledged effects reconciledChange and recover data safely
On a realistic table copy, build a customer-history index concurrently outside a transaction block and check its catalog validity. Hold a conflicting transaction open, run a nullable-column migration under a short session-local lock budget, and show it fails promptly. Resolve the blocker through its owner, then complete the compatible migration. Generate old row versions and demonstrate how an idle transaction affects cleanup. Detach a disposable old partition, restore its archive in another database, and only then mark it eligible for deletion. Finally inject an incorrect bulk update, recover a fresh instance to the chosen marker, and compare records immediately before and after the target. Fence the original writer and reconcile any externally acknowledged payouts after that marker.
Cost and verification
Record disk growth, additional WAL and archive storage, connector replay time, index-build I/O, migration lock wait, and measured restore time. The drill passes only when a synthetic payout can be read from the recovered application path, no business effect was duplicated, and the last proven recovery point is stated precisely. A running database process or an advancing connector offset is insufficient evidence. Remove the isolated resources after the evidence is retained.
Common Mistakes
- Do not drop a lagging slot without an offset recovery decision.
- Do not run a concurrent index build inside a migration transaction.
- Do not accept a PITR target before reconciling external effects.
Connected lessons
- Replication slots: bound retained WAL before a consumer outage fills the disk
- CDC snapshot handoff: prove that initial rows and later changes form one history
- Transactional outbox: commit business state and event intent together
- Concurrent index builds: verify validity after the command exits
- DDL lock budget: fail a migration before it stalls production traffic
- Vacuum horizon: find transactions holding old row versions alive
- Partition retention: detach, prove the archive, then delete
- Point-in-time recovery: accept a restored timeline only after business reconciliation
- DevOps projects
