A backtest must replay what the detector knew at each decision and distinguish unreviewed alerts from confirmed non-incidents.
Rare-event backtests: delayed labels, windows and incident recall
Replay by decision time
For each hour, rebuild the baseline from earlier complete periods only, apply source-health state known then, and use the threshold selected on an older development window. A random split of hourly rows lets later seasonal behavior and corrected labels leak backward. Rolling-origin evaluation gives a more honest sequence of decisions.
Judge incidents as episodes
Match alerts to confirmed incident windows under a written tolerance, such as first detection within 47 minutes of onset. Count each incident once even when it generates many low-count hours. Report lead time, missed incidents, duplicate pages and investigation volume; pointwise accuracy is dominated by the many normal hours and tells little about response value.
Handle incomplete labels
A recently fired alert may still be under review. Mark it pending and exclude it from a confirmed precision denominator while still showing its count. Mature only those decisions whose label-delay window has elapsed. The label contract keeps outcome date separate from event date, preventing silent survivorship bias.
Test late confirmation
Create an alert on day 6 and a confirmed incident label on day 9. A report run on day 7 should show the alert pending; a report run on day 10 can credit the detection. Add an uninvestigated alert from day 6 and ensure it does not become a false positive merely because no label arrived.
Implementation
def mature_alert_labels(alert_rows, report_day, review_delay_days):
mature = []
pending = []
for alert in alert_rows:
if alert["day"] + review_delay_days <= report_day and alert["reviewed"]:
mature.append(alert)
else:
pending.append(alert)
return mature, pendingPerformance and operating cost
A pass over A alerts costs O(A) time and O(A) output space. Episode matching can cost more when windows overlap; index incidents by tenant and time, and preserve the matching rule with every reported metric.
Common Mistakes
- Do not calculate accuracy over overwhelmingly normal hours as the sole result.
- Do not call pending or uninvestigated alerts false positives.
- Do not use a threshold or baseline fitted on the held-out incident window.
