A query and a relevant document can use different languages. Preserve intent, protected identifiers and document provenance through retrieval.
Cross-language retrieval: query, document and locale contracts
Separate user language from evidence language
A support engineer may ask in Hindi for the rollback limit while the runbook is written in English. Record query language, requested answer language, document language and service scope separately. Do not infer that a shared script means a shared language, or that an English identifier should be translated. Keep the raw query and its normalized form tied to one request ID. Code-switching policy matters when one sentence combines a local-language question with an English service name.
Protect operational terms
Service IDs, error codes, release tags and command names must survive translation or cross-language embedding. A translated phrase can be fluent yet point to the wrong service. Represent protected terms as typed slots and validate that every candidate query retains them. Keep a glossary version for domain terms whose ordinary-language meaning differs from their operational meaning. Translation placeholders provide the same safeguard when producing a localized reply.
Keep multiple retrieval routes inspectable
Candidate generation may combine a translated query, a multilingual embedding and exact matches for protected identifiers. Store which route produced each candidate and the version of its index or model. Access filters must run before a passage is exposed to ranking or answer generation. A translated query is a proposal, not a replacement for the original user intent. If language detection is uncertain, search under several plausible language hypotheses and preserve that uncertainty for review.
Test actual cross-language relevance
Create query-document judgments where the query and evidence use different languages. Include false friends, untranslated product names, mixed-script tickets and documents whose wording changed between revisions. Measure recall before ranking and final relevance by language pair and service. A high monolingual score does not establish cross-language quality. The ranking audit verifies that a retrieved passage supports the requested answer rather than merely sharing a translated keyword.
Implementation
def prepare_retrieval_request(raw_query, query_language, answer_language,
protected_terms, translated_candidates):
if not raw_query.strip() or not query_language or not answer_language:
raise ValueError("query and language codes are required")
accepted = [candidate for candidate in translated_candidates
if all(term in candidate for term in protected_terms)]
return {"raw_query": raw_query, "query_language": query_language,
"answer_language": answer_language,
"protected_terms": tuple(sorted(protected_terms)),
"candidate_queries": accepted,
"state": "ready" if accepted else "review"}
request = prepare_retrieval_request(
"gateway-west ka rollback limit kya hai?", "hi", "hi",
{"gateway-west"},
["What is the rollback limit for gateway-west?",
"What is the rollback limit for the gateway?"])
assert request["candidate_queries"] == [
"What is the rollback limit for gateway-west?"]
assert request["state"] == "ready"
Performance and operating cost
Validating c candidates against p protected terms takes O(c × p × n) in the worst case for candidate length n with simple substring checks. Translation, embedding and index search have their own costs. The example is an admission check, not a translation engine; production matching should use typed slots and boundary-aware validation so one identifier cannot accidentally match inside another.
Common Mistakes
- Translating a service ID as ordinary prose.
- Dropping the original query after translation.
- Evaluating only same-language query-document pairs.
- Ranking restricted passages before access policy is applied.
Read next
- Cross-language ranking and source-language evidence review
- Project: answer a local-language question from a foreign-language runbook
- Script profiles and code-switching boundaries in text intake
- Translation contracts for placeholders, numbers and terminology
- Hybrid text ranking with access filters and reranking
Continue the workflow: Cross-script linkage: contextual evidence and collision review.
