Create a language-aware router that uses transfer where it helps and hands off cases without enough native-language evidence.
Project: adapt a support-intent router with scarce local labels
Define the release decision
A support team has reviewed examples for three language slices and a small new-locale queue. The router may assign a known intent only when its language-specific audit clears a floor; otherwise it creates a manual-review item with the original message intact. An inferred intent never authorizes a payment or account action. Keep a reason code so agents know whether the handoff came from low confidence, scant evidence or a disputed label.
Build the corpus
Group conversations and customers before splitting. Keep the final audit native, recent and untouched by generation. Annotate ambiguous terms with local reviewers and record disagreements. Compare a lexical baseline, a shared multilingual model and a small adapted model under one label policy. The transfer boundary sets the support floor; augmentation controls prevent misleading gains.
Choose thresholds by queue cost
For each language and critical intent, estimate wrong-route rate against the manual-review volume produced by a threshold. Include rare billing and access cases even when their support count is low. Shadow-route traffic before automating; sample both confident and rejected cases for review. A model that saves agent time on common requests may still be unacceptable if it misroutes a rare high-cost request.
Ship and monitor
Version model, tokenizer, label map and thresholds together. Show original text to authorized reviewers, not a translation-only substitute. Monitor intent precision, reject rate, false automation and queue age by language. On drift, fall back to manual routing for the affected slice rather than rolling back every language. Retain deletion and privacy controls for both original and generated training examples.
Implementation
def propose_intent_route(case, audit_floor, threshold_by_language):
language = case["language"]
if case["reviewed_examples"] < audit_floor:
return {"route": "manual", "reason": "audit-floor"}
threshold = threshold_by_language.get(language)
if threshold is None or case["score"] < threshold:
return {"route": "manual", "reason": "uncertain"}
return {"route": case["intent"], "reason": "reviewed-slice"}
sample = {"language": "locale-k", "reviewed_examples": 23,
"score": 0.86, "intent": "order-status"}
assert propose_intent_route(sample, 20, {"locale-k": 0.82})["route"] == "order-status"
Performance and operating cost
The gate is O(1) per case; model inference and human handoff dominate runtime and operating cost. Track minutes spent per manual case, not just model request latency. One shared model may be economical, but a per-language audit and reject policy are still required before it can route customer work.
Common Mistakes
- Letting the model act on an account after classifying an intent.
- Using translated audit text in place of native messages.
- Setting one threshold for every language from pooled accuracy.
- Counting synthetic variants as independent reviewed cases.
Read next
- Low-resource NLP: label budgets and transfer boundaries
- Low-resource augmentation: provenance, label drift and slice audit
- Calibrate multilingual text decisions and fallback routes
- Project: route support tickets with auditable text evaluation
- Dialogue handoff: gate actions on verified state and uncertainty
