Rank foreign-language passages against a query while keeping a verifiable path back to the original source text.
Cross-language ranking and source-language evidence review
Judge the source passage
A translated snippet may be easier to read, but the evidence is the original passage at a particular revision. Store its source-language span, document ID, access level and retrieval score. A localized display can accompany the passage, yet it must not silently replace the source text during verification. Passage revision auditing prevents an answer from relying on a stale translation of an updated runbook.
Separate recall from ranking
If the relevant foreign-language document never enters the candidate set, reranking cannot recover it. Measure candidate recall at the chosen depth, then rank quality over the same judgments. Compare translated-query, multilingual-vector and exact-ID routes by language pair. Merge candidates by document and passage identity; do not mistake two translations of one passage for two independent sources. Hybrid ranking supplies a useful baseline, but the cross-language judgment set must contain meaning-based relevance rather than keyword overlap alone.
Verify support before answering
The top passage can mention the right service while contradicting the query premise. Check the requested environment, time, version and action against the original span. Record whether the answer is directly supported, requires multiple passages or remains unanswered. If no reviewer can verify the source language, route the result to a qualified reviewer or abstain. Evidence contracts prevent a fluent localized answer from becoming unsupported advice.
Report uneven performance
Aggregate scores can hide failure for a low-volume language pair. Report candidate recall, rank positions, verified answer support and abstention by query language, document language, service and script pattern. Include code-switched queries and translated terms with several domain meanings. The multilingual search project gates answers on source passage identity and language-aware review, not merely on a cosine score.
Implementation
def first_relevant_rank(ranked_passage_ids, judged_relevant_ids):
for rank, passage_id in enumerate(ranked_passage_ids, start=1):
if passage_id in judged_relevant_ids:
return rank
return None
def reciprocal_rank_by_pair(evaluations):
totals = {}
for case in evaluations:
language_pair = (case["query_language"], case["document_language"])
rank = first_relevant_rank(case["ranked_passages"],
set(case["relevant_passages"]))
score, count = totals.get(language_pair, (0.0, 0))
totals[language_pair] = (score + (1 / rank if rank else 0.0), count + 1)
return {pair: score / count for pair, (score, count) in totals.items()}
cases = [{"query_language": "hi", "document_language": "en",
"ranked_passages": ["gateway-47", "runbook-82"],
"relevant_passages": ["runbook-82"]}]
assert reciprocal_rank_by_pair(cases) == {("hi", "en"): 0.5}
assert first_relevant_rank(["gateway-47"], {"runbook-82"}) is None
Performance and operating cost
For q cases with k ranked passages each, scoring takes O(q × k) time and O(q) space in the worst case for per-pair aggregates, excluding relevance-set creation. Reciprocal rank measures placement of the first judged relevant passage, not whether an answer is supported or whether every source language has adequate coverage. Keep those checks separate.
Common Mistakes
- Treating a localized snippet as the canonical evidence.
- Reporting reranking quality without candidate recall.
- Counting two translations of one passage as independent support.
- Accepting a top hit that names the service but misses the environment or revision.
