Skip to content
AITroveRead. Build. Understand.
Make this comfortable

RAG evaluation: separate retrieval recall from answer support

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A retrieval-based answer can fail because evidence was missed, misused or unavailable; evaluate each boundary with different tests.

Build judged questions

Write questions from real user intents with allowed corpus revision, required passage IDs and expected answer or no-answer state. Include exact identifiers, paraphrases, conflicting policies and queries outside the corpus. Keep an untouched final set after tuning chunking, retrieval and prompts. Document versions] make an evaluation item reproducible.

Score retrieval first

Check whether all required evidence fits inside the selected context budget. A relevant passage at rank 40 is effectively absent when only six passages reach generation. Report recall by query type and access scope. If evidence is missing, changing answer wording will not fix the root cause. Retrieval choices] should be tested against the same judged set.

Score answers separately

When supporting passages were present, inspect factual claims against those passages, scope, no-answer behavior and conflicting-evidence handling. An answer can be fluent yet unsupported. Automated graders can prioritize review, but sample manual judgments and disagreement checks are needed for high-impact claims. Track evidence IDs and answer status in the result.

Keep an error ledger

For each failure, classify missing corpus content, permission filter, chunk boundary, rank, context truncation, generation misuse or unsupported assertion. This assigns repairs to the right stage. Re-run the same frozen questions after one change at a time; otherwise a score movement cannot be attributed.

Implementation

python
def retrieval_recall_at_budget(required_ids, ranked_ids, budget):
    if budget <= 0:
        raise ValueError("context budget must be positive")
    required = set(required_ids)
    if not required:
        raise ValueError("use a no-answer test for zero required passages")
    returned = set(ranked_ids[:budget])
    return len(required & returned) / len(required)

case_result = retrieval_recall_at_budget(judged_passage_ids, ranked_passage_ids, 6)

Performance and operating cost

Checking one case uses O(R+B) time for R required IDs and B returned IDs. Building high-quality judged sets costs expert review; that cost is necessary to know which stage changed.

Common Mistakes

  • Do not use answer fluency as a retrieval metric.
  • Do not count a passage beyond the context budget as retrieved evidence.
  • Do not tune on the final evaluation set repeatedly.

Read next

ai-data
retrieval-ai
Storage details