Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Anomaly triage: deduplicate, explain and close the feedback loop

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

The useful output is an investigation packet with enough context to act, a stable incident key and a recorded disposition.

Group related alerts

Group by tenant, metric and overlapping time window before paging. Keep the strongest score and the first-seen time; attach later observations to the same open incident. A grouping key that is too broad can hide separate outages, while one that includes every hour creates a new ticket each time. Alert capacity should be measured after this grouping.

Show the evidence

Include observed count, baseline, comparable periods, source-health status, exposure and relevant release marker. Explain what the score measured, not just that a model emitted 0.93. The packet should reveal whether a tenant has only a few historical comparison periods and whether a missing upstream partition changes the interpretation.

Record disposition

Let an investigator mark confirmed incident, expected behavior, source failure or unresolved, with event IDs and a review timestamp. Retain reason and label delay. Feed reviewed outcomes into the next threshold review only after a frozen evaluation run; immediately training on operator feedback can bias the detector toward cases that happened to be investigated.

Test replay and reopening

Send three consecutive hourly alerts with the same incident key, then replay the middle event. The ticket count should remain one. Resolve it, then send a new candidate after the configured quiet interval and check that a new incident opens with a distinct ID. A source-failure disposition should reach source owners without altering a business baseline.

Implementation

python
def new_incident_alerts(alert_rows):
    earliest_by_key = {}
    for alert in alert_rows:
        key = alert["incident_key"]
        if key not in earliest_by_key or alert["observed_at"] < earliest_by_key[key]["observed_at"]:
            earliest_by_key[key] = alert
    return list(earliest_by_key.values())

Performance and operating cost

Deduplicating A alerts is O(A) expected time and O(I) memory for I incident keys. A production key needs a time-bounded state store and a reopen policy; indefinitely retaining every key would grow without limit.

Common Mistakes

  • Do not page on each observation in one continuing incident.
  • Do not train immediately on selectively reviewed tickets.
  • Do not hide the raw observation and source status behind a score.

Read next

ai-data
anomaly-detection
Storage details