Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Correction-aware quality alerts

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A correction-aware alert evaluates a published interval and its revision history without counting the same late change as a fresh incident on every replay.

Anchor the interval and version

An account-day total may be published at noon and corrected after a late settlement arrives. Store event-time interval, source cutoff, table generation, correction generation and quality result. An alert key based only on calendar day conflates distinct evidence; a key based only on run ID creates a new incident for every retry. Late-arrival policy defines how long an interval may legitimately change.

Separate provisional from final checks

A window still accepting late events can have an expected shortfall. Apply a provisional threshold before closure and a stricter final reconciliation after the correction horizon. The status should say provisional, final or reopened. A consumer reading a provisional number needs that state attached to the release, especially when the serving index lags the warehouse.

Deduplicate alerts by cause

If the first failed run, its automatic retry and a manual replay all see the same missing source shard, create one incident with attempt evidence attached. A later independent loss after the source recovers can open another incident. Use the dataset, interval and failure signature as a stable key; close the incident only when a committed generation passes the check. Run evidence identifies which attempts were published.

Prevent false recovery

A late correction can improve a metric while still leaving the row count below the source control. Require all release gates to pass before resolving the incident. Check totals, unique keys, source positions and the consumer-visible index generation. A green task and a quiet alert channel do not prove the old dashboard answer was replaced.

Exercise replay and reopen

Publish an interval with 47 missing transactions, retry it twice, then apply a 31-transaction correction. The incident should remain open because 16 are still absent. Apply the final 16 and publish one corrected generation; now resolve it. Replaying the identical correction must not reopen or double-count. If a later source gap appears, its new failure signature should create a separate case.

Implementation

python
control_count = 470
published = {"generation": 71, "count": 423}
corrections = [(72, 31), (73, 16), (73, 16)]

def apply_unique_corrections(base, changes):
    count = base["count"]
    seen = set()
    for generation, increment in changes:
        if generation in seen:
            continue
        seen.add(generation)
        count += increment
    return count

assert apply_unique_corrections(published, corrections[:1]) == 454
assert apply_unique_corrections(published, corrections) == control_count

Performance and operating cost

The replay scan is O(C) expected time and O(C) deduplication state for C correction records. Production systems can use a stable correction ID and bounded retention, but that horizon must cover supported replays. Incident state adds one record per active dataset-interval-signature combination; excessive partition-level keys can overwhelm operators and should be rolled up by a meaningful failure boundary.

Common Mistakes

  • Do not create one alert per retry of the same source failure.
  • Do not close an incident when only one of several release checks recovered.
  • Do not apply a duplicate late correction twice.

Read next

Continue the workflow: Contract gate severity and release evidence.

ai-data
data-engineering
Storage details