A cutover gate compares old and new representations at a fixed input position and refuses publication when correctness, coverage or replay guarantees differ.
Migration reconciliation and cutover gates
Choose a stable comparison point
A migrating order ledger continues to receive events. Comparing old output at position 81 with new output at position 84 creates a false mismatch. Pin both calculations to the same source position, schema revision and event-time interval. Store the pin in a manifest. For a stream, wait until both paths have acknowledged the position; wall-clock completion alone is not a consistency boundary.
Compare more than totals
Equal aggregate revenue can hide one duplicate positive amount and one missing negative adjustment. Check primary-key sets, duplicate counts, row-level amount hashes, null rates, currency distribution and totals by region and date. Treat accepted exclusions as named rules with counts. Quality quarantine gives mismatches a place to wait without contaminating a published reader.
Separate candidate from visible state
Write the new representation and index as candidate generations. Keep readers on the last accepted pointer until every check passes. If an index build fails after the table is correct, the release remains incomplete. An atomic pointer switch must identify the table and serving index generations together; otherwise the application can query a new table through an old index.
Prove rollback and delayed replay
Before cutover, replay a delayed event encoded with the old schema and verify the new path still interprets it. Then move the reader pointer, force a failed post-cutover check, and restore the old pointer. Preserve the new candidate for diagnosis. Rollback of the pointer does not reverse external side effects, so downstream exports need their own versioned handoff or pause window.
Make failure observable
A failed gate should identify the interval, input position, first differing key, affected count and failed predicate without exposing sensitive row payloads in public logs. Record gate duration separately from migration throughput. A pass at one small canary interval is not proof for the full history; run the same controls across every partition included in the release.
Implementation
old_rows = {"order-47": 2375, "order-48": -125, "order-49": 6400}
new_rows = {"order-47": 2375, "order-48": -125, "order-49": 6400}
def reconciliation_report(previous, candidate):
keys = previous.keys() | candidate.keys()
mismatched = [key for key in keys if previous.get(key) != candidate.get(key)]
return {"equal": not mismatched, "mismatched_keys": sorted(mismatched),
"previous_total": sum(previous.values()), "candidate_total": sum(candidate.values())}
assert reconciliation_report(old_rows, new_rows)["equal"]
new_rows["order-48"] = -100
assert reconciliation_report(old_rows, new_rows)["mismatched_keys"] == ["order-48"]Performance and operating cost
The reference comparison is O(N) expected time and O(N) key state for N records. A distributed version can compare partition digests first, then inspect only mismatched partitions, but digest equality alone is not a substitute for control totals and key counts. Candidate generations temporarily double storage, and each gate adds reads before publication.
Common Mistakes
- Do not compare outputs produced from different source positions.
- Do not accept equal totals as proof of equal records.
- Do not switch the table pointer before the serving index is ready.
Read next
- Expand-contract schema migration
- Data quality gates: quarantine bad rows and reconcile complete batches
- Table snapshots and atomic publication
- Serving indexes and freshness contracts
- Project: migrate a live payment amount contract
Continue the workflow: Cross-engine shadow reads and discrepancy triage.
