New words and changed message length are warning signals. They do not prove the router’s decisions have become wrong.
Text drift: input shifts, delayed labels and real errors
Separate input from outcome
Monitor observable changes at intake: language mix, token lengths, unknown terms, channel distribution and abstention. Actual routing errors require reviewed labels, which may arrive days later. A keyword spike can reflect a product launch while model precision stays steady; a stable vocabulary can hide a changed meaning of “hold.” Keep distinct dashboards for input shift, delayed quality and user impact. Temporal validation gives a baseline for comparison.
Avoid logging raw private text
Count approved aggregate features and sample restricted messages for human review only under policy. Do not put customer names, order IDs or full transcripts into ordinary metrics. Keep windows, model version, tokenizer version and source channel explicit. If a classifier begins rejecting more messages, investigate whether the mix changed or the threshold bundle changed. Serving contracts should carry those versions.
Account for label delay
An agent’s final queue may be an imperfect proxy for the correct intent, and it is not available at prediction time. Define when a reviewed outcome is mature enough for comparison. Use event time for intake cohorts and avoid mixing recent unlabeled cases with older fully reviewed ones. Report sample counts and pending outcomes. A model must not train on its own unverified routes as if they were independent truth.
Trigger investigation, not blind retraining
Set alerts on sustained shifts and quality ceilings, then inspect representative errors by slice. Retrain only after the label policy, data provenance and audit show a real need. A rollback may be safer when a new model or tokenizer caused the change. Feedback controls cover learning from outcomes; the project runs the monitoring workflow.
Implementation
def drift_window(window, minimum_count, unknown_rate_ceiling):
if window["count"] < minimum_count:
return "insufficient-volume"
if window["unknown_term_rate"] > unknown_rate_ceiling:
return "investigate-input-shift"
if window["reviewed_count"] < minimum_count:
return "await-reviewed-outcomes"
return "inspect-quality-metrics"
recent = {"count": 240, "unknown_term_rate": 0.19,
"reviewed_count": 42}
assert drift_window(recent, 100, 0.15) == "investigate-input-shift"
Performance and operating cost
The per-window gate is O(1); aggregating n requests is O(n) time and bounded metric storage when labels and windows are fixed. Reviewed labels and root-cause analysis dominate cost. Do not retrain merely because an inexpensive signal moved; compare against mature outcomes and the cost of a wrong route.
Common Mistakes
- Calling a vocabulary change proof of prediction failure.
- Comparing recent unlabeled traffic with older reviewed traffic as equals.
- Logging raw customer messages into monitoring metrics.
- Training on the model’s own routes as ground truth.
Read next
- Text feedback loops: correction provenance and shadow audits
- Project: monitor support-text drift without a feedback trap
- Text validation: split conversations, duplicates and time together
- Text inference: package tokenizer, labels and reject paths
- Text classification evaluation: inspect slices and allow abstention
Continue the workflow: Log event parameters: typed identity, drift and privacy.
