Ship a support router that keeps mixed-language input intact, measures each reviewed slice and sends uncertain requests to a staffed queue.
Project: route multilingual tickets with measured abstention
Define the decision first
The endpoint chooses returns, billing or delivery for a support ticket. It may also return manual-review with a reason. A case does not disappear when language is unknown. Capture the original message, conversation identity, arrival time, channel and reviewed queue. Add a separate reviewed language and script-mix label for evaluation; never infer a gold language from the model’s own prediction. Record who can see raw ticket text and where it is retained.
Prepare a hard audit set
Keep conversations, copied replies and repeated product IDs in one split. Reserve recent messages as the final audit. Within that set, ask reviewers to identify Romanized Hindi, mixed Devanagari and English, pasted English signatures, emoji-heavy short tickets and unsupported languages. The audit needs enough examples in each slice to support a claim; if it does not, report uncertainty and route that slice for review. Use script profiling only as an observable input characteristic.
Compare candidate systems
Train a sparse baseline and a multilingual model against the same split and taxonomy. Compare per-queue recall, wrong-queue cost, automatic coverage and reviewed-slice errors. A more complex model earns deployment only if it improves the operational decision after its latency and review costs are counted. Tune thresholds on validation data, then freeze them before final audit. Follow calibration and fallback rather than selecting a single score that flatters the average.
Release and keep watching
The serving bundle must identify tokenizer, model, label order, thresholds and policy version. For every decision, store an audit record without exposing raw text to general logs. Review a random sample of confident automatic routes and all abstentions; otherwise correction data is selected by the old policy. Track shift in traffic composition and staff queue time. If a slice degrades, lower automatic coverage for that slice and roll back the bundle when needed. Connect this project to the first ticket-routing project and the serving contract.
Implementation
def audit_routing(decisions):
by_slice = {}
for decision in decisions:
reviewed_slice = decision["reviewed_slice"]
stats = by_slice.setdefault(reviewed_slice, {"total": 0, "automatic": 0, "correct": 0})
stats["total"] += 1
if decision["reason"] == "automatic":
stats["automatic"] += 1
stats["correct"] += decision["predicted_queue"] == decision["reviewed_queue"]
return {reviewed_slice: {
"coverage": stats["automatic"] / stats["total"],
"selective_accuracy": stats["correct"] / stats["automatic"] if stats["automatic"] else None,
"reviewed_cases": stats["total"],
} for reviewed_slice, stats in by_slice.items()}
audit_rows = [
{"reviewed_slice": "mixed", "reason": "automatic", "predicted_queue": "returns", "reviewed_queue": "returns"},
{"reviewed_slice": "mixed", "reason": "below-threshold", "predicted_queue": None, "reviewed_queue": "billing"},
]
assert audit_routing(audit_rows)["mixed"]["coverage"] == 0.5
Performance and operating cost
The audit pass is O(n) time and O(s) space for n reviewed decisions and s slices. Model inference and staff review dominate live cost; report p95 latency per script mix and expected review cases per hour. Selective accuracy excludes abstentions by definition, so publish coverage next to it. A slice with no automatic decisions has undefined selective accuracy, not a perfect score.
Common Mistakes
- Using model-predicted language as the evaluation ground truth.
- Reporting selective accuracy without coverage.
- Allowing automated routing for an unsupported slice because its score is high.
- Learning only from abstained tickets and missing confident errors.
Read next
- Script profiles and code-switching boundaries in text intake
- Calibrate multilingual text decisions and fallback routes
- Project: route support tickets with auditable text evaluation
- Text inference: package tokenizer, labels and reject paths
- Project: ship an auditable support-entity extractor
Continue the workflow: Project: audit support routing across language varieties.
