The useful output is an investigation packet with enough context to act, a stable incident key and a recorded disposition.
Anomaly triage: deduplicate, explain and close the feedback loop
Group related alerts
Group by tenant, metric and overlapping time window before paging. Keep the strongest score and the first-seen time; attach later observations to the same open incident. A grouping key that is too broad can hide separate outages, while one that includes every hour creates a new ticket each time. Alert capacity should be measured after this grouping.
Show the evidence
Include observed count, baseline, comparable periods, source-health status, exposure and relevant release marker. Explain what the score measured, not just that a model emitted 0.93. The packet should reveal whether a tenant has only a few historical comparison periods and whether a missing upstream partition changes the interpretation.
Record disposition
Let an investigator mark confirmed incident, expected behavior, source failure or unresolved, with event IDs and a review timestamp. Retain reason and label delay. Feed reviewed outcomes into the next threshold review only after a frozen evaluation run; immediately training on operator feedback can bias the detector toward cases that happened to be investigated.
Test replay and reopening
Send three consecutive hourly alerts with the same incident key, then replay the middle event. The ticket count should remain one. Resolve it, then send a new candidate after the configured quiet interval and check that a new incident opens with a distinct ID. A source-failure disposition should reach source owners without altering a business baseline.
Implementation
def new_incident_alerts(alert_rows):
earliest_by_key = {}
for alert in alert_rows:
key = alert["incident_key"]
if key not in earliest_by_key or alert["observed_at"] < earliest_by_key[key]["observed_at"]:
earliest_by_key[key] = alert
return list(earliest_by_key.values())Performance and operating cost
Deduplicating A alerts is O(A) expected time and O(I) memory for I incident keys. A production key needs a time-bounded state store and a reopen policy; indefinitely retaining every key would grow without limit.
Common Mistakes
- Do not page on each observation in one continuing incident.
- Do not train immediately on selectively reviewed tickets.
- Do not hide the raw observation and source status behind a score.
