Detect input shifts, wait for mature reviewed outcomes and compare a shadow router before changing production.
Project: monitor support-text drift without a feedback trap
Set the monitoring boundary
The production router emits a proposal, reject reason and model bundle version. Aggregate intake language, channel, length and unknown-term signals without logging unrestricted text. A restricted audit store receives a governed sample for independent review. The alerting service may request investigation, but it cannot automatically retrain and deploy from one shifted metric.
Build a credible baseline
Choose stable historical windows with mature reviewed labels and record traffic mix. Include launch periods, channel changes, copied templates and rare high-cost intents. Group related tickets during evaluation. Input and delayed-label separation prevents a recent unlabeled week from looking falsely healthy; feedback controls preserve independent evidence.
Investigate and shadow
When unknown terms or abstention rise, inspect restricted sampled cases and determine whether a product term, tokenizer or user intent changed. Compare a candidate model in shadow mode on the same intake cohort. Wait for independent reviewed outcomes before deciding. Report false routes, reject rate, reviewer time and queue age by relevant slice. A new model must pass the fixed audit and current traffic audit.
Release or roll back
Version text preprocessing, model, taxonomy and thresholds together. Promote only after review of high-cost failures; keep a prior bundle for rollback. If traffic changes again, recheck the same metrics rather than assuming a one-time fix. Retention deletion must propagate to raw audit samples and derived training rows, while aggregate counts remain policy-compliant.
Implementation
def candidate_release_decision(baseline, shadow, minimum_reviews):
if shadow["reviewed_count"] < minimum_reviews:
return "wait-for-labels"
if shadow["wrong_route_rate"] > baseline["wrong_route_rate"]:
return "hold-release"
if shadow["review_minutes"] > baseline["review_minutes"]:
return "investigate-workload"
return "eligible-for-review"
baseline = {"wrong_route_rate": 0.08, "review_minutes": 470}
shadow = {"reviewed_count": 180, "wrong_route_rate": 0.06,
"review_minutes": 430}
assert candidate_release_decision(baseline, shadow, 120) == "eligible-for-review"
Performance and operating cost
The decision gate is O(1); shadow inference and independent labels create temporary extra cost. The calculation does not replace statistical analysis or risk review. Compare mature cohorts and high-cost mistakes before promoting a model, because a cheaper queue can hide more rework elsewhere.
Common Mistakes
- Automatically retraining after a keyword spike.
- Comparing shadow and production on different traffic cohorts.
- Using agent silence as proof that a route was correct.
- Keeping deleted tickets in a shadow training copy.
