Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Frozen baselines and alert investigation

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A monitoring alert is reproducible only when the baseline, extraction, measurement contract and response rule are frozen before the new period arrives.

Separate baseline and monitoring phases

Estimate the late-response center and chart parameters from a designated historical window. Then freeze a versioned baseline for later days. If an outlying day is folded back into the center as soon as it appears, the chart adapts to the very change it was meant to detect. Rebaseline only after the cause is understood and the process definition has intentionally changed. A deployment that alters response timestamps may require a new measurement contract before any statistical comparison.

Triage the data path first

For a signaled day, verify eligible case count, event timestamp coverage, deduplication and query version. Compare raw case IDs and numerator IDs with the daily aggregate. Check whether yesterday’s data were complete at the extraction cutoff. A logging outage can produce a low rate that looks like excellent performance, while a missing denominator can produce an artificial spike. Reconciliation should precede stories about staff behavior.

Investigate operations second

If the data path is intact, inspect queue mix, site, shift, campaign timing, staffing and incident records. Keep contemporaneous evidence rather than explaining the alert from memory a week later. A chart says that the current metric is unusual under its baseline model; it does not isolate which factor moved. Do not split into many subgroups until one produces an attractive narrative without accounting for repeated searching.

Record four distinct outcomes

The day may miss the service target, trigger a process signal, do both, or do neither. For each alert, record the observed rate, case count, versioned baseline limits, operational target, cause hypothesis, checks performed and final disposition. A confirmed data error requires corrected aggregates and replay of downstream EWMA state. A confirmed process shift may require intervention and later a reviewed new baseline. Stateful monitoring makes backfill handling especially important.

Design a release gate

Do not send a public chart when its eligible count is zero, its baseline version is missing, or the current query version is incompatible with the baseline. An alert can still be raised internally as an instrumentation incident. Store a resolved or unresolved disposition for each signal, including who reviewed it and when. The project implements a compact packet builder and a release decision.

Implementation

python
def alert_packet(day, late_cases, eligible_cases, query_version,
                     baseline_version, baseline_query_version,
                     observed_rate, lower, upper, service_target):
    if eligible_cases <= 0 or not 0 <= late_cases <= eligible_cases:
        raise ValueError("invalid daily counts")
    if abs(observed_rate - late_cases / eligible_cases) > 1e-12:
        raise ValueError("reported rate does not reconcile")
    if not baseline_version or not query_version:
        raise ValueError("missing provenance")
    if not 0 <= lower <= upper <= 1 or not 0 <= service_target <= 1:
        raise ValueError("invalid monitoring thresholds")
    compatible = query_version == baseline_query_version
    return {"day": day, "baseline_version": baseline_version,
            "eligible_cases": eligible_cases, "late_cases": late_cases,
            "target_miss": observed_rate > service_target,
            "process_signal": observed_rate < lower or observed_rate > upper,
            "publishable": compatible}

packet = alert_packet("shift-47", 21, 140, "response-v3", "base-2026q3",
                      "response-v3", 0.15, 0.0, 0.112, 0.03)
assert packet["process_signal"] and packet["target_miss"]
assert packet["publishable"]

Performance and operating cost

A packet for one day is O(1) time and space once daily counts and limits exist. Reconciling N underlying case IDs costs O(N) expected time with hashed IDs. Incident review time grows with evidence collection and should not be hidden inside a chart computation.

Common Mistakes

  • Do not rebaseline automatically after an inconvenient signal.
  • Do not investigate operations before verifying denominator and timestamp integrity.
  • Do not publish a chart across incompatible metric versions.

Read next

ai-data
data-science
Storage details