A detector should report when its input contract fails before interpreting that failure as a change in real-world behavior.
Anomaly operations: distinguish drift, incidents and broken inputs
Check source health first
Watch event ingestion lag, missing partitions, deduplication rate and schema errors alongside the business metric. A quiet receipt stream caused by a failed collector needs an ingestion alert, not a claim that customers stopped buying. Stream parity checks can identify when the online count diverges from a corrected batch view.
Separate time scales
An incident may be a sudden local departure; drift is a sustained change in the baseline or population. A planned rollout can move volume without being either a failure or a healthy permanent shift. Maintain an event log of releases and calendar changes, and require a deliberate baseline update after investigation instead of allowing every new value to teach the detector immediately.
Track detector health
Plot scores, alert rate, pending-label fraction and baseline age by tenant. If one segment stops producing alerts while its source becomes stale, that is not improved accuracy. Monitor fallbacks for newly active tenants and changes in exposure quality. Model monitoring should join outcome data only when it matures.
Rehearse three scenarios
Simulate a genuine 80 percent receipt drop with healthy ingestion, an ingestion halt with unknown business volume, and a gradual planned expansion. The first should enter incident triage, the second source-health triage, and the third a baseline-review workflow. Record which evidence made each route possible at decision time.
Implementation
def classify_volume_signal(source_healthy, observed, baseline, planned_change):
if not source_healthy or observed is None:
return "source-health"
if planned_change:
return "baseline-review"
if observed < baseline * 0.2:
return "incident-candidate"
return "ordinary"Performance and operating cost
Each classification is O(1) time and space after baseline lookup. Health metrics need their own timely collection path; a detector cannot reliably diagnose a broken source using only the source it is watching.
Common Mistakes
- Do not translate a stale source into a business anomaly.
- Do not absorb an active incident into the baseline automatically.
- Do not assume a falling alert rate proves performance improved.
