Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Evaluate translation by adequacy, terminology and locale slices

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Reference overlap can screen translation changes, but a release needs human checks for meaning, protected fields and harm in each supported locale.

Freeze comparable inputs

Construct a reviewed audit set from actual supported workflows, not only clean benchmark sentences. Keep near-duplicate conversations in one split. Record source, reference translation, locale, domain, protected tokens and reviewer notes. A single source may have several valid translations, so an overlap score can penalize acceptable wording. Use corpus-level automated metrics as a regression signal and retain their tokenization settings with the report. The split contract avoids copied ticket leakage.

Review meaning and obligation

Ask bilingual reviewers whether the target preserves actor, action, tense, negation, amount and uncertainty. Score terminology separately from fluency. A translation that sounds native yet says a refund has arrived when the source says it is pending is operationally wrong. Review source-to-target and target-to-source where helpful, but back-translation alone is not proof because two errors can cancel. Protected-field validation catches deterministic failures first.

Report slices honestly

Break out short requests, long instructions, mixed scripts, formal and informal registers, product-specific vocabulary and each locale. Small slices should show support counts and uncertainty, not a confident green badge. Track the rate of untranslatable or ambiguous source messages and the number sent for clarification. If a locale lacks reviewed evidence, keep human approval as the default. The script profile may describe input, but it is not a locale truth label.

Set a release gate

Require zero protected-token failures on the audit set, a minimum reviewed adequacy level by supported locale, an acceptable harmful-error count and a documented human fallback. Compare the same inputs across old and new bundles. A model update can improve average overlap while regressing a rare safety instruction. Monitor production correction reasons and human edit time after release. The reply project applies the gate.

Implementation

python
def translation_release_gate(locale_audits):
    if not locale_audits:
        raise ValueError("at least one reviewed locale is required")
    for locale, audit in locale_audits.items():
        if audit["reviewed_cases"] < 47:
            return False, f"insufficient-review:{locale}"
        if audit["protected_token_failures"] or audit["harmful_meaning_errors"]:
            return False, f"critical-error:{locale}"
        if audit["adequacy_pass_rate"] < 0.91:
            return False, f"adequacy:{locale}"
    return True, "reviewed-locales-pass"

audits = {"hi-IN": {"reviewed_cases": 53, "protected_token_failures": 0,
                     "harmful_meaning_errors": 0, "adequacy_pass_rate": 0.94}}
assert translation_release_gate(audits)[0]

Performance and operating cost

The gate is O(l) time and O(1) extra space for l locale audits. Human bilingual evaluation is the main cost; sample by risk and traffic, but keep a fixed minimum for every locale you claim to support. Automated metrics are cheap once references exist, yet they can miss semantic reversals. Report reviewer time, rejection reasons and post-release edits with model latency.

Common Mistakes

  • Using one aggregate overlap score as proof of translation quality.
  • Calling a locale supported with only a handful of reviewed cases.
  • Treating protected-token success as evidence that meaning is preserved.
  • Changing metric tokenization between model versions.

Read next

ai-data
natural-language-processing
Storage details