Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: gate a payment mart on quality evidence

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Publish a payment mart only when cohort drift, source completeness and late-correction evidence agree with the release contract.

Freeze the source controls

Create one payment-day interval with 470 eligible attempts and a control count for each region. Keep the source snapshot and schema ID in the manifest. The candidate mart initially contains 423 records because one shard was delayed. A global success rate may still look plausible; cohort baselines and input coverage must be checked separately.

Build a provisional candidate

Calculate per-region success rates from deduplicated payment attempts, label the interval provisional and keep it off the final reader pointer. Record numerator, denominator, cohort size and reference window. A cohort below its minimum sample size yields insufficient evidence, not pass. Allow a dated last-good mart to serve while the candidate is quarantined.

Replay corrections once

Deliver 31 delayed records, retry the same correction message, then deliver 16 more. Use correction IDs or source positions so the retry has no effect. The incident remains open after the first correction and closes only after the source count reaches 470 and regional totals reconcile. Keep all candidate attempts in the run manifest.

Commit and check the consumer

Publish a new table snapshot and its serving index as one release generation. Query payment success rate and count through the consumer route; verify both refer to that generation. Run a rollback simulation that restores the previous pointer without deleting evidence. A correction that passes warehouse checks but never reaches the index is not a successful release.

Submit the evidence package

Include source controls, candidate and final row counts, cohort baseline version, two correction IDs, duplicate replay result, incident timeline, committed snapshot and query output. Add a deliberately missing shard test and a legitimate high-volume promotion test. They should not be confused merely because each changes the raw transaction count.

Implementation

python
source_control = {"north": 230, "south": 240}
candidate = {"north": 230, "south": 193}
correction_batches = [("batch-72", "south", 31), ("batch-73", "south", 16),
                      ("batch-72", "south", 31)]

def reconcile(rows, changes, expected):
    result, seen = dict(rows), set()
    for correction_id, region, count in changes:
        if correction_id not in seen:
            result[region] += count
            seen.add(correction_id)
    return result == expected

assert not reconcile(candidate, correction_batches[:1], source_control)
assert reconcile(candidate, correction_batches, source_control)

Performance and operating cost

Reconciliation is O(R + C) expected time for R regions and C correction messages, with O(R + C) state. The project also retains candidate snapshots and a last-good serving generation; these add storage but make rejected releases inspectable. A full regional baseline scan costs more than one global aggregate, so schedule it at the release boundary rather than on every query.

Common Mistakes

  • Do not pass a cohort check when its source denominator is incomplete.
  • Do not treat a duplicate correction as new revenue.
  • Do not publish a corrected table while the serving index still points at stale data.

Read next

Continue the workflow: Project: ship a payment feed with executable contract gates.

ai-data
data-engineering
Storage details