Keyword and embedding results solve different failures. Combine them only after access filtering, then rerank candidates against the actual query.
Hybrid text ranking with access filters and reranking
Start with the authorized corpus
A search service can rank a private incident note perfectly and still be wrong if the requester cannot read it. Apply tenant and document permissions before candidates reach the user or a generation model. If an index cannot enforce prefiltering, verify every candidate against a current authorization source before returning it, and refill from a deeper candidate pool. Postfiltering can reduce recall; log that separately. Never let a cached answer bypass a changed access policy.
Keep exact and semantic signals
Lexical retrieval handles literal error codes, product names and rare identifiers. Dense retrieval can connect paraphrases, such as “charged twice” with a passage about duplicate capture. Each can fail alone. Fetch a bounded list from both, deduplicate by document revision and passage ID, and combine rank positions rather than raw scores whose scales differ. A reciprocal-rank formula is a straightforward baseline. The rank constant and candidate depth are tuning parameters, not universal numbers. Embedding pairs determine whether the dense side has learned the task.
Rerank a small candidate set
A query-passage reranker reads both together and can resolve finer relations than a precomputed vector. Its cost grows with candidate count and text length, so invoke it only on the authorized merged list. Preserve original scores and the reason each passage entered the list for failure review. Truncate carefully: removing the sentence that contains the answer can make a correct candidate look irrelevant. Keep a lexical fallback when the embedding service or reranker times out.
Evaluate the whole decision
Use reviewed queries with relevant document IDs, allowed user identities and a declared “no answer” state. Report recall at candidate depth before reranking, rank of first useful hit after reranking, access violations, p95 latency and fallback rate. A retrieval change that lifts rank while exposing one unauthorized passage is a failed release. The runbook search project makes these checks a deployment gate.
Implementation
def reciprocal_rank_fusion(lexical_ids, semantic_ids, allowed_ids, rank_constant=47):
if rank_constant <= 0:
raise ValueError("rank constant must be positive")
fused = {}
for candidate_list in (lexical_ids, semantic_ids):
seen_in_list = set()
for rank, document_id in enumerate(candidate_list, start=1):
if document_id not in allowed_ids or document_id in seen_in_list:
continue
seen_in_list.add(document_id)
fused[document_id] = fused.get(document_id, 0.0) + 1 / (rank_constant + rank)
return sorted(fused, key=lambda document_id: (-fused[document_id], document_id))
ranked = reciprocal_rank_fusion(["public-47", "private-19"],
["public-47", "public-82"], {"public-47", "public-82"})
assert ranked == ["public-47", "public-82"]
Performance and operating cost
For l lexical and d dense candidate IDs, fusion is O(l + d + u log u) time for u unique authorized candidates and O(u) space. A cross-encoder reranker adds roughly one model pass per candidate, so cap its list and measure tail latency. Permission checks may dominate if performed one document at a time; batch them without weakening current-user authorization. Refresh authorization on cached results when memberships change.
Common Mistakes
- Adding raw lexical and cosine scores without calibration.
- Filtering private passages only after they have entered a generated answer.
- Treating a high-ranked passage as proof that an answer exists.
- Ignoring the drop in recall caused by permission filtering or truncation.
Read next
- Text embeddings: pair labels, hard negatives and versioned vectors
- Project: release an incident-runbook search service
- Retrieval-Based AI Tutorial
- Text inference: package tokenizer, labels and reject paths
- Text validation: split conversations, duplicates and time together
Continue the workflow: Project: release an incident-runbook search service.
Continue the workflow: Project: release morphology-aware support search.
Continue the workflow: Long documents: section boundaries and passage provenance.
Continue the workflow: Query rewrites: provenance, protected tokens and user intent.
Continue the workflow: Cross-language ranking and source-language evidence review.
Continue the workflow: Retrieval negatives: hard candidates and false-negative risk.
