A fictional parcel service is releasing order notices and a support assistant in English, Hindi, and French Canada. Build a release packet that contains source strings, locale tags, glossary version, machine placeholders, typed monetary and date values, account-permission checks, current policy evidence, and matched evaluation cases. The prompt may draft language and identify ambiguous requests. It cannot rewrite order IDs, decide a customer's account scope, or format a date by guessing the user's timezone. The final artifact is a release decision with exact failures and owners.
Project: release a multilingual order notice safely
Prepare three workflows
First, localize the shipment notice for OR-958 while preserving shipment_id and dispatch_date. Run the placeholder check, then ask a French Canada reviewer to inspect meaning and terminology. Second, render the amount 1,234.50 CAD and dispatch instant 2026-11-04T18:30:00Z from typed fields after resolving locale and time zone. Third, route 'Refund kab milega for OR-958? Please answer in English' to a permission-checked refund-status read. A quoted Spanish sentence asking what it means must route to translation, not a private order lookup.
Resolve policy and evaluate
Retrieve the current policy PL-62 and older localized page PL-41 for a delayed parcel question. Keep both evidence IDs and effective dates. Explain that the current rule concerns store credit, flag the ambiguous older wording, and do not issue any effect without the account and policy checks. Build matched English and Hindi cases with a missing refund record: both must abstain. Include a French notice with placeholders and plural or formatting checks, plus the quoted Spanish control. If one locale invents a refund date, block promotion even if the average fluency rating improves.
OR-958: preserve order ID, shipment_id, dispatch_date.
Typed amount=1234.50 CAD; instant=2026-11-04T18:30:00Z.
Mixed Hindi/English request: refund-status read, English reply.
Spanish quoted text: translation only; no account lookup.
Current PL-62 beats expired PL-41; flag wording conflict.
Gate: equal permitted actions and abstention across matched locales.Performance and operating cost
For C cases, L locales, and V prompt variants, comparison uses O(CLV) model evaluations. Placeholder scans are linear in string length, and formatter work is linear in typed fields for a fixed message. Human linguistic review costs more than a token check but catches semantic shifts the check cannot see. Record policy, glossary, formatter, and prompt versions so a failed locale can be reproduced. Separate failure counts for wrong intent, missing placeholders, wrong format, unsupported policy claim, and unauthorized effect; one broad quality score will conceal these differences.
Common Mistakes
- Do not treat a fluent translation as proof that placeholders or policy meaning survived.
- Do not let language detection authorize a private data read.
- Do not average away a wrong action in one supported locale.
Connected lessons
- Translation prompts: preserve terms and machine placeholders
- Locale-aware prompts: keep values typed until rendering
- Mixed-language requests: separate intent from reply language
- Cross-language retrieval: keep policy meaning attached to its source
- Cross-locale evaluation: compare decisions, not word-for-word text
- Prompt templates: bind variables without changing instruction structure
- Tool calls: validate intent and arguments before an external effect
- HTML multilingual case page project: preserve names and a clear reading order
- Multilingual prompt workflow decisions
