A retrieval stage must find relevant permitted passages before generation; different query types need different matching signals.
Retrieval design: compare lexical, semantic and reranked candidates
Start from query failures
A user may ask about a precise invoice code or a paraphrased policy. Exact terms favor lexical search; paraphrases may benefit from semantic similarity. Combine candidates under a fixed budget and rerank them using the actual query. Do not assume one embedding score represents answer correctness. Chunk design] changes what either retriever can find.
Filter before ranking output
Apply tenant and document permissions at retrieval time. Deduplicate overlapping chunks and cap the number of passages from one parent document. Preserve passage IDs and revisions in the ranked result. A reranker that receives forbidden text has already crossed the access boundary, even if the final response hides it.
Measure retrieval directly
Build questions with judged supporting passage IDs, including exact-code, paraphrase, multi-section and unanswerable cases. Report recall at the context budget and rank position of the first required passage. A generation failure cannot be diagnosed if retrieval evidence was never measured. Layered evaluation] separates these stages.
Control latency
Retrieve a broader candidate set cheaply, then rerank a smaller set. Measure lexical, vector and reranking time separately. A higher recall may not justify doubling p95 latency if the answer task rarely needs the added passages; decide under a stated service budget.
Implementation
def combine_candidates(lexical_hits, semantic_hits, allowed_ids, limit=12):
if limit <= 0:
raise ValueError("candidate limit must be positive")
scores, passages = {}, {}
for hits in (lexical_hits, semantic_hits):
for rank, passage in enumerate(hits, start=1):
if passage.chunk_id not in allowed_ids:
continue
scores[passage.chunk_id] = scores.get(passage.chunk_id, 0) + 1 / (60 + rank)
passages[passage.chunk_id] = passage
ranked_ids = sorted(scores, key=scores.__getitem__, reverse=True)
selected, per_document = [], {}
for chunk_id in ranked_ids:
passage = passages[chunk_id]
if per_document.get(passage.document_id, 0) >= 2:
continue
selected.append(passage)
per_document[passage.document_id] = per_document.get(passage.document_id, 0) + 1
if len(selected) == limit:
break
return selectedPerformance and operating cost
Merging L lexical and S semantic hits uses O(L+S) expected hash work, then O(K log K) sorting for K unique hits. Reranking adds query-passage model cost per candidate.
Common Mistakes
- Do not infer answer quality from vector similarity alone.
- Do not expose forbidden text to a reranker.
- Do not skip retrieval recall checks and blame generation for missing evidence.
Read next
- Chunking documents: preserve section identity and answer boundaries
- RAG evaluation: separate retrieval recall from answer support
- Grounded answers: require passage support and a no-answer path
- Retrieval corpus contracts: identity, permissions and document lifecycle
Continue the workflow: Candidate retrieval: separate broad discovery from hard eligibility.
Continue the workflow: Ranking query groups and relevance labels.
