A correction-aware alert evaluates a published interval and its revision history without counting the same late change as a fresh incident on every replay.
Correction-aware quality alerts
Anchor the interval and version
An account-day total may be published at noon and corrected after a late settlement arrives. Store event-time interval, source cutoff, table generation, correction generation and quality result. An alert key based only on calendar day conflates distinct evidence; a key based only on run ID creates a new incident for every retry. Late-arrival policy defines how long an interval may legitimately change.
Separate provisional from final checks
A window still accepting late events can have an expected shortfall. Apply a provisional threshold before closure and a stricter final reconciliation after the correction horizon. The status should say provisional, final or reopened. A consumer reading a provisional number needs that state attached to the release, especially when the serving index lags the warehouse.
Deduplicate alerts by cause
If the first failed run, its automatic retry and a manual replay all see the same missing source shard, create one incident with attempt evidence attached. A later independent loss after the source recovers can open another incident. Use the dataset, interval and failure signature as a stable key; close the incident only when a committed generation passes the check. Run evidence identifies which attempts were published.
Prevent false recovery
A late correction can improve a metric while still leaving the row count below the source control. Require all release gates to pass before resolving the incident. Check totals, unique keys, source positions and the consumer-visible index generation. A green task and a quiet alert channel do not prove the old dashboard answer was replaced.
Exercise replay and reopen
Publish an interval with 47 missing transactions, retry it twice, then apply a 31-transaction correction. The incident should remain open because 16 are still absent. Apply the final 16 and publish one corrected generation; now resolve it. Replaying the identical correction must not reopen or double-count. If a later source gap appears, its new failure signature should create a separate case.
Implementation
control_count = 470
published = {"generation": 71, "count": 423}
corrections = [(72, 31), (73, 16), (73, 16)]
def apply_unique_corrections(base, changes):
count = base["count"]
seen = set()
for generation, increment in changes:
if generation in seen:
continue
seen.add(generation)
count += increment
return count
assert apply_unique_corrections(published, corrections[:1]) == 454
assert apply_unique_corrections(published, corrections) == control_countPerformance and operating cost
The replay scan is O(C) expected time and O(C) deduplication state for C correction records. Production systems can use a stable correction ID and bounded retention, but that horizon must cover supported replays. Incident state adds one record per active dataset-interval-signature combination; excessive partition-level keys can overwhelm operators and should be rolled up by a meaningful failure boundary.
Common Mistakes
- Do not create one alert per retry of the same source failure.
- Do not close an incident when only one of several release checks recovered.
- Do not apply a duplicate late correction twice.
Read next
- Watermark retention and late corrections
- Run-scoped lineage and release evidence
- Cohort baselines and denominator drift
- Project: gate a payment mart on quality evidence
- Data quality gates: quarantine bad rows and reconcile complete batches
Continue the workflow: Contract gate severity and release evidence.
