A production model alert is useful only when it identifies customer impact, an owner and an action that can be taken now.
Model alerts: page on customer symptoms with a named owner
Separate immediate symptoms from slow evidence
An endpoint timing out is visible now; model accuracy may take 23 days to measure because outcomes arrive later. Page on acute customer symptoms such as timeout rate, invalid responses or a broken fallback. Send feature drift and uncertain quality changes to investigation unless they cross a predeclared operational boundary. A shifted histogram alone does not establish that customers received wrong decisions. Delayed-label monitoring keeps these clocks separate, while latency observation captures the live request path.
Attach ownership and a first action
For each alert, name the on-call team, affected service, model digest, comparison window, expected response and rollback or mitigation route. A page that says only “feature distribution changed” forces the responder to invent the question during an incident. Put a dashboard link in the operational alert configuration, but keep raw personal data out of broad notifications. Route source schema failures to the producer owner and model-quality changes to the model owner, with one coordinator when both are involved.
Control noise with counts and duration
A ratio from four requests can look dramatic. Require a minimum population, a sustained breach or both, while preserving a fast path for severe safety failures. Group related symptoms under one incident rather than paging separately for queue growth, timeout and fallback volume when they share a cause. Use low-cardinality labels such as service and release; customer ID labels can overwhelm metric storage. A muted alert must have an expiry and owner so it does not become a permanent blind spot.
Test the alert as a product
Inject a feature-store outage, a model crash, slow requests and a delayed-label gap. Verify that the right owner receives one actionable page for customer impact and that nonurgent monitoring remains visible without repeated pages. Measure detection delay, false pages and time to identify the release. The incident drill requires a responder to use the alert record, not tribal knowledge, to choose mitigation.
Implementation
def alert_decision(window):
total = window["requests"]
if total < 470:
return {"state": "observe", "reason": "small-window"}
timeout_rate = window["timeouts"] / total
if window["unsafe_decisions"] > 0:
return {"state": "page", "reason": "unsafe-decision"}
if timeout_rate > 0.025 and window["breach_minutes"] >= 7:
return {"state": "page", "reason": "sustained-timeouts"}
if window["feature_drift"]:
return {"state": "investigate", "reason": "drift-signal"}
return {"state": "healthy"}
window = {"requests": 500, "timeouts": 19, "unsafe_decisions": 0,
"breach_minutes": 8, "feature_drift": False}
assert alert_decision(window)["state"] == "page"
assert alert_decision({**window, "timeouts": 0,
"feature_drift": True})["state"] == "investigate"
Performance and operating cost
A window decision is O(1) time and space once counts are aggregated. Histograms and counters cost storage proportional to label combinations and retention. The example uses illustrative limits; production thresholds need traffic baselines, sampling checks and a policy for severe low-volume failures.
Common Mistakes
- Paging on every drift statistic as if it proved customer harm.
- Alerting on a ratio without a denominator.
- Sending a page with no owner or mitigation step.
- Muting a noisy alert indefinitely instead of repairing its rule.
Read next
- Model incidents: build a release and evidence timeline before rollback
- Project: run a receipt-model incident drill with honest mitigation
- Model monitoring: separate input drift, data faults and delayed outcomes
- Inference latency budgets: measure queue, feature and model time
- Serving overload: bound queues and choose a fallback before time runs out
Continue the workflow: Model monitoring dimensions without metric-cardinality failure.
