A translation can read naturally while changing an order identifier, amount, safety instruction or variable placeholder. Protect nonlinguistic tokens before evaluating style.
Translation contracts for placeholders, numbers and terminology
Classify what may change
A customer message may contain a currency amount, date, order ID, product name and a template marker such as {ticket_id}. Some values require exact preservation; others need locale-aware formatting. Write a per-field policy. A placeholder should be protected by a typed sentinel whose mapping is stored securely, then restored and validated after translation. Do not assume the model will copy braces or digits faithfully. Keep the source revision, target locale and glossary version in the request manifest.
Separate language quality from field integrity
A fluent sentence with the wrong amount is a failed translation. Run deterministic checks for protected tokens, expected identifiers, markup and numeric policy before human quality review. A repeated identifier may occur twice; compare multiplicity, not just set membership. Allow grammatical inflection of approved terms where the locale requires it, and review terminology in context. Unicode handling matters when normalization changes visible punctuation or digits.
Make ambiguity visible
A short source may not specify who acted, whether a refund is completed or what “it” refers to. Translators and models should not invent that context. Retain an uncertainty marker and request clarification for customer-facing high-impact copy. Preserve negation and modality. Review source and target side by side, including protected fields and back-linked evidence. The evaluation path covers adequacy and locale slices beyond exact-token checks.
Package the serving rule
Version model, tokenizer, glossary, locale normalization, sentinel encoding and restore validator as one bundle. A failure to restore all placeholders must stop publication. Include inputs with adjacent placeholders, mixed scripts, repeated order IDs and right-to-left text in the audit set. The localized-reply project uses this contract before a message reaches a customer.
Implementation
from collections import Counter
import re
PLACEHOLDER = re.compile(r"\{[a-z][a-z0-9_]*\}")
def validate_placeholders(source_text, translated_text):
source = Counter(PLACEHOLDER.findall(source_text))
translated = Counter(PLACEHOLDER.findall(translated_text))
if source != translated:
raise ValueError("placeholder names or counts changed")
return translated
source = "Case {ticket_id} is assigned to {agent_name}."
target = "{agent_name} is assigned case {ticket_id}."
assert validate_placeholders(source, target)["{ticket_id}"] == 1
Performance and operating cost
Scanning source and target is O(n + m) time and O(p) space for p distinct placeholders. Translation inference and bilingual review dominate operating cost. Deterministic validation is cheap and should run on every result, but it cannot prove meaning preservation. A protected token may survive while the sentence reverses negation, so combine field checks with reviewed adequacy samples.
Common Mistakes
- Checking placeholder sets but missing a duplicated or deleted occurrence.
- Assuming all numbers should be copied unchanged across locale formats.
- Accepting fluent text that reverses negation or completion status.
- Changing glossary or sentinel rules without versioning the translation bundle.
Read next
- Evaluate translation by adequacy, terminology and locale slices
- Project: release reviewed localized support replies
- Unicode and tokenization: preserve meaning at the text boundary
- Calibrate multilingual text decisions and fallback routes
- Text inference: package tokenizer, labels and reject paths
Continue the workflow: Evaluate translation by adequacy, terminology and locale slices.
Continue the workflow: Grammar correction: edit spans without changing protected meaning.
