Publish a payment mart only when cohort drift, source completeness and late-correction evidence agree with the release contract.
Project: gate a payment mart on quality evidence
Freeze the source controls
Create one payment-day interval with 470 eligible attempts and a control count for each region. Keep the source snapshot and schema ID in the manifest. The candidate mart initially contains 423 records because one shard was delayed. A global success rate may still look plausible; cohort baselines and input coverage must be checked separately.
Build a provisional candidate
Calculate per-region success rates from deduplicated payment attempts, label the interval provisional and keep it off the final reader pointer. Record numerator, denominator, cohort size and reference window. A cohort below its minimum sample size yields insufficient evidence, not pass. Allow a dated last-good mart to serve while the candidate is quarantined.
Replay corrections once
Deliver 31 delayed records, retry the same correction message, then deliver 16 more. Use correction IDs or source positions so the retry has no effect. The incident remains open after the first correction and closes only after the source count reaches 470 and regional totals reconcile. Keep all candidate attempts in the run manifest.
Commit and check the consumer
Publish a new table snapshot and its serving index as one release generation. Query payment success rate and count through the consumer route; verify both refer to that generation. Run a rollback simulation that restores the previous pointer without deleting evidence. A correction that passes warehouse checks but never reaches the index is not a successful release.
Submit the evidence package
Include source controls, candidate and final row counts, cohort baseline version, two correction IDs, duplicate replay result, incident timeline, committed snapshot and query output. Add a deliberately missing shard test and a legitimate high-volume promotion test. They should not be confused merely because each changes the raw transaction count.
Implementation
source_control = {"north": 230, "south": 240}
candidate = {"north": 230, "south": 193}
correction_batches = [("batch-72", "south", 31), ("batch-73", "south", 16),
("batch-72", "south", 31)]
def reconcile(rows, changes, expected):
result, seen = dict(rows), set()
for correction_id, region, count in changes:
if correction_id not in seen:
result[region] += count
seen.add(correction_id)
return result == expected
assert not reconcile(candidate, correction_batches[:1], source_control)
assert reconcile(candidate, correction_batches, source_control)Performance and operating cost
Reconciliation is O(R + C) expected time for R regions and C correction messages, with O(R + C) state. The project also retains candidate snapshots and a last-good serving generation; these add storage but make rejected releases inspectable. A full regional baseline scan costs more than one global aggregate, so schedule it at the release boundary rather than on every query.
Common Mistakes
- Do not pass a cohort check when its source denominator is incomplete.
- Do not treat a duplicate correction as new revenue.
- Do not publish a corrected table while the serving index still points at stale data.
Read next
- Cohort baselines and denominator drift
- Correction-aware quality alerts
- Data quality gates: quarantine bad rows and reconcile complete batches
- Serving indexes and freshness contracts
- Replay manifests and audit trails
Continue the workflow: Project: ship a payment feed with executable contract gates.
