A cross-locale evaluation should pair semantically equivalent requests across supported languages while allowing natural differences in expression. Score whether each locale reaches the same permitted action, preserves required entities and amounts, handles uncertainty, and obeys the same refusal boundary. Exact string matching is usually a poor measure of translation quality. Include cases with mixed language, missing data, plural forms, quoted text, and policy version conflicts. Keep localized gold labels and reviewer notes versioned; a policy change can invalidate a past expected answer without making the old run disappear.
Cross-locale evaluation: compare decisions, not word-for-word text
Operational case
An order assistant is evaluated on a matched English and Hindi request for refund status on OR-958, a French Canada dispatch notice containing two placeholders, and a Spanish quotation that asks for a translation rather than an order lookup. The English variant correctly abstains when the refund record is unavailable, but the Hindi variant invents an expected date. That mismatch blocks release even if the French notice reads well. Reviewers log whether the failure came from intent extraction, missing evidence, or answer generation, then rerun both languages on the same account and policy snapshots after the prompt changes.
Case pair OR-958: English and Hindi; same missing refund record.
Expected action: abstain in both; preserve OR-958.
Notice fr-CA: placeholder counts match; amount/date renderer checked.
Spanish quote: translate text; no order lookup.
Release: block if any locale invents a private status or date.Performance and operating cost
With C semantic cases, L locales, and V prompt variants, a full comparison needs O(CLV) model evaluations plus human language review. Sampling can reduce routine workload, but every high-impact action boundary should retain a matched case for each supported locale. Separate translation fluency from decision correctness and placeholder integrity so an average score does not hide a dangerous route in one language. Store prompt, model, glossary, policy, and formatter versions for each run. A reviewer should be able to explain why two phrasings are equivalent without requiring identical output wording.
Common Mistakes
- Do not grade translated answers solely by exact string equality.
- Do not allow a failure in one language to be averaged away by another.
- Do not compare locale runs made against different policy snapshots.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Multilingual prompts: test policy meaning across languages
- Evaluation sets: measure the failure cases that matter
- Paired prompt evaluation: count changes, then inspect uncertainty
- Translation prompts: preserve terms and machine placeholders
- Locale-aware prompts: keep values typed until rendering
- Mixed-language requests: separate intent from reply language
- Cross-language retrieval: keep policy meaning attached to its source
- Project: release a multilingual order notice safely
- Multilingual prompt workflow decisions
Continue with: Localization prompts: use pseudolocale and real layout checks.
