Skip to content
AITroveRead. Build. Understand.
Make this comfortable

CDC snapshot handoff: prove that initial rows and later changes form one history

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

Change-data capture often begins with a table snapshot and then reads the database log from a recorded position. The two phases must cover changes made while the snapshot runs without a gap. A restart before snapshot completion may cause rows to be emitted again; a failover can change which log positions remain available. The downstream sink therefore needs stable keys and a repeatable upsert or deduplication rule.

Operational decision

A claims search index is rebuilt from PostgreSQL while new claims keep arriving. Record connector configuration, table set, snapshot completion marker, source log position, and sink checkpoint before starting. Apply snapshot records by claim ID and version, then stream updates from the connector's supported handoff point; do not start the stream at a guessed wall-clock time. Inject an update and a delete while the snapshot scans, then stop the connector immediately before it records completion. On restart, verify the final search index matches a bounded source query and that no claim is missing or resurrected. For a primary failover, compare the connector's saved offset with the new source's slot position before allowing it to publish more events. If the required history is unavailable, rebuild from a new snapshot rather than silently skipping to the current log. Preserve an audit record of source row count, sink row count, and sampled content hashes at the acceptance boundary.

Output
Claims CDC handoff record
Source tables and schema revision: captured
Snapshot: completed marker and source log position
Sink key: claim ID plus monotonic row version
Restart test: update and delete during snapshot; no gap or resurrection
Failover gate: saved offset is available on new source
Acceptance: bounded source and sink reconciliation matches

Cost and verification

A consistent snapshot can consume I/O and retain old row versions while it runs. Deduplicating writes adds sink storage and CPU but makes replay safe. Large table scans should be paced against user latency and replication pressure. A count match can hide wrong rows, so compare identifiers and sampled values too. An offset advancing without a sink checkpoint is not proof that downstream data is durable.

Common Mistakes

  • Do not choose a CDC start point by wall-clock guesswork.
  • Do not assume a restart will never repeat snapshot rows.
  • Do not advance past an unavailable offset without an explicit rebuild plan.

Connected lessons

Practice and check

Kafka operating follow-up

devops
operations
Storage details