Create a governed quality review for support-intent routing, with explicit disagreement and a safe fallback for unsupported slices.
Project: audit support routing across language varieties
Set the scope
The support router predicts operational intent, not a writer’s identity or education. A separate restricted audit samples messages across language varieties that the service actually receives. Reviewers label request intent and note disputed interpretations under one policy. Serving does not persist a dialect profile. Uncertain cases go to an authorized agent rather than a forced queue.
Build the reviewed set
Select consented or appropriately governed examples, group conversation families and preserve original wording. Include local abbreviations, code-switching, negation and messages with identical intent expressed differently. Record reviewer disagreements before adjudication. Slice design avoids proxy-identity claims; the gate controls sample support and privacy.
Compare policy changes
Measure per-intent precision and recall, false routing, abstention and agent time by supported slice. Compare a baseline and revised model on the same frozen audit. Inspect whether a threshold change simply shifts errors into manual review. The team should see severe mistakes with their source context, not only averages, while broad reports suppress tiny groups.
Operate and retire data
Version model, taxonomy and slice definition. Review drift after channel or tokenizer changes. Keep raw audit messages under limited access and retention; delete derived examples when required. If an audit slice fails the release criterion, hold the affected locale or channel rollout until a revised model or policy passes review. Do not silently fall back to a high-resource language prediction.
Implementation
def quality_route(prediction, supported_locales, thresholds):
locale = prediction.get("declared_locale")
if locale not in supported_locales:
return {"state": "agent-review", "reason": "insufficient-audit"}
if prediction["score"] < thresholds[locale]:
return {"state": "agent-review", "reason": "uncertain-intent"}
return {"state": "proposed", "intent": prediction["intent"]}
case = {"declared_locale": "locale-k", "score": 0.84,
"intent": "delivery"}
assert quality_route(case, {"locale-k"}, {"locale-k": 0.81})["state"] == "proposed"
Performance and operating cost
The routing gate is O(1) per case and uses a declared locale; detailed linguistic slices remain in the restricted audit. Annotation, adjudication and agent review dominate cost. Track wrong routes and manual load together before enabling more automation.
Common Mistakes
- Turning a quality-audit slice into a permanent customer trait.
- Using translated messages as the only evidence for local writing.
- Suppressing small-slice uncertainty to claim a release pass.
- Removing the agent fallback because aggregate accuracy rose.
Read next
- Dialect-aware NLP audits: slices and label disagreement
- Dialect audit release gates and privacy boundaries
- Project: route multilingual tickets with measured abstention
- Low-resource NLP: label budgets and transfer boundaries
- Adjudicate text labels before they become training truth
Continue the workflow: Project: moderate a support thread with quotes and appeal review.
