Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Rare-event backtests: delayed labels, windows and incident recall

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A backtest must replay what the detector knew at each decision and distinguish unreviewed alerts from confirmed non-incidents.

Replay by decision time

For each hour, rebuild the baseline from earlier complete periods only, apply source-health state known then, and use the threshold selected on an older development window. A random split of hourly rows lets later seasonal behavior and corrected labels leak backward. Rolling-origin evaluation gives a more honest sequence of decisions.

Judge incidents as episodes

Match alerts to confirmed incident windows under a written tolerance, such as first detection within 47 minutes of onset. Count each incident once even when it generates many low-count hours. Report lead time, missed incidents, duplicate pages and investigation volume; pointwise accuracy is dominated by the many normal hours and tells little about response value.

Handle incomplete labels

A recently fired alert may still be under review. Mark it pending and exclude it from a confirmed precision denominator while still showing its count. Mature only those decisions whose label-delay window has elapsed. The label contract keeps outcome date separate from event date, preventing silent survivorship bias.

Test late confirmation

Create an alert on day 6 and a confirmed incident label on day 9. A report run on day 7 should show the alert pending; a report run on day 10 can credit the detection. Add an uninvestigated alert from day 6 and ensure it does not become a false positive merely because no label arrived.

Implementation

python
def mature_alert_labels(alert_rows, report_day, review_delay_days):
    mature = []
    pending = []
    for alert in alert_rows:
        if alert["day"] + review_delay_days <= report_day and alert["reviewed"]:
            mature.append(alert)
        else:
            pending.append(alert)
    return mature, pending

Performance and operating cost

A pass over A alerts costs O(A) time and O(A) output space. Episode matching can cost more when windows overlap; index incidents by tenant and time, and preserve the matching rule with every reported metric.

Common Mistakes

  • Do not calculate accuracy over overwhelmingly normal hours as the sole result.
  • Do not call pending or uninvestigated alerts false positives.
  • Do not use a threshold or baseline fitted on the held-out incident window.

Read next

ai-data
anomaly-detection
Storage details