Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Anomaly thresholds: rank evidence within an alert budget

Last updated: 7 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

The threshold is an operating decision: it trades missed incidents against investigation load and repeated pages.

Set a capacity constraint

If a team can investigate 12 new tenant incidents a day, an alert rule producing 90 tickets is unusable even if its point-level recall looks high. Define the daily budget, response hours and severity tiers before choosing a percentile. A critical source outage can bypass the ordinary queue, but that exception needs a separate rule and owner.

Rank with interpretable evidence

Combine contextual residual, source health, duration and affected tenant count in a score whose components can be inspected. A lower count is not always worse: five receipts against an expectation of eight differs from five against 460. The contextual baseline gives the score its meaning. Calibrate the threshold on a historical period that predates the held-out test.

Measure the whole queue

Report alert volume, unique incident count, confirmation fraction, time to first alert and missed confirmed incidents. Show the distribution by tenant size and region so a single noisy group does not consume the queue. The threshold should be reviewed when workload changes, but do not retune it against the final test period after seeing the outcome.

Exercise a fixed budget

Given candidate severities 0.94, 0.72, 0.66 and 0.31 with capacity two, the selected alerts are the first two after applying eligibility and deduplication. If both top scores belong to the same ongoing incident, the second investigation slot should remain available to the next distinct incident.

Implementation

python
def select_alerts(candidates, capacity):
    if capacity < 0:
        raise ValueError("capacity cannot be negative")
    if capacity == 0:
        return []
    ranked = sorted(candidates, key=lambda item: item["severity"], reverse=True)
    selected = []
    seen_incidents = set()
    for candidate in ranked:
        incident_key = candidate["incident_key"]
        if incident_key not in seen_incidents:
            seen_incidents.add(incident_key)
            selected.append(candidate)
        if len(selected) == capacity:
            break
    return selected

Performance and operating cost

Sorting N candidates costs O(N log N) time and O(N) memory. Streaming top-K selection can reduce retained candidates to O(K), but deduplication and severity updates still need an incident-state policy.

Common Mistakes

  • Do not choose a threshold using the final held-out period.
  • Do not confuse point-level precision with distinct-ticket workload.
  • Do not page repeatedly for one unresolved incident.

Read next

ai-data
anomaly-detection
Storage details