Reference overlap can screen translation changes, but a release needs human checks for meaning, protected fields and harm in each supported locale.
Evaluate translation by adequacy, terminology and locale slices
Freeze comparable inputs
Construct a reviewed audit set from actual supported workflows, not only clean benchmark sentences. Keep near-duplicate conversations in one split. Record source, reference translation, locale, domain, protected tokens and reviewer notes. A single source may have several valid translations, so an overlap score can penalize acceptable wording. Use corpus-level automated metrics as a regression signal and retain their tokenization settings with the report. The split contract avoids copied ticket leakage.
Review meaning and obligation
Ask bilingual reviewers whether the target preserves actor, action, tense, negation, amount and uncertainty. Score terminology separately from fluency. A translation that sounds native yet says a refund has arrived when the source says it is pending is operationally wrong. Review source-to-target and target-to-source where helpful, but back-translation alone is not proof because two errors can cancel. Protected-field validation catches deterministic failures first.
Report slices honestly
Break out short requests, long instructions, mixed scripts, formal and informal registers, product-specific vocabulary and each locale. Small slices should show support counts and uncertainty, not a confident green badge. Track the rate of untranslatable or ambiguous source messages and the number sent for clarification. If a locale lacks reviewed evidence, keep human approval as the default. The script profile may describe input, but it is not a locale truth label.
Set a release gate
Require zero protected-token failures on the audit set, a minimum reviewed adequacy level by supported locale, an acceptable harmful-error count and a documented human fallback. Compare the same inputs across old and new bundles. A model update can improve average overlap while regressing a rare safety instruction. Monitor production correction reasons and human edit time after release. The reply project applies the gate.
Implementation
def translation_release_gate(locale_audits):
if not locale_audits:
raise ValueError("at least one reviewed locale is required")
for locale, audit in locale_audits.items():
if audit["reviewed_cases"] < 47:
return False, f"insufficient-review:{locale}"
if audit["protected_token_failures"] or audit["harmful_meaning_errors"]:
return False, f"critical-error:{locale}"
if audit["adequacy_pass_rate"] < 0.91:
return False, f"adequacy:{locale}"
return True, "reviewed-locales-pass"
audits = {"hi-IN": {"reviewed_cases": 53, "protected_token_failures": 0,
"harmful_meaning_errors": 0, "adequacy_pass_rate": 0.94}}
assert translation_release_gate(audits)[0]
Performance and operating cost
The gate is O(l) time and O(1) extra space for l locale audits. Human bilingual evaluation is the main cost; sample by risk and traffic, but keep a fixed minimum for every locale you claim to support. Automated metrics are cheap once references exist, yet they can miss semantic reversals. Report reviewer time, rejection reasons and post-release edits with model latency.
Common Mistakes
- Using one aggregate overlap score as proof of translation quality.
- Calling a locale supported with only a handful of reviewed cases.
- Treating protected-token success as evidence that meaning is preserved.
- Changing metric tokenization between model versions.
Read next
- Translation contracts for placeholders, numbers and terminology
- Project: release reviewed localized support replies
- Text validation: split conversations, duplicates and time together
- Script profiles and code-switching boundaries in text intake
- Text classification evaluation: inspect slices and allow abstention
