Search relevance evaluation compares a fixed query set, a known catalog snapshot, eligible candidates, human judgments, and observed ranking. Beacon Supply has 47 fictional help and catalog queries. Five return no eligible result and 42 return at least one. A model can draft intent labels and discrepancy notes, while a reviewer validates evidence and a search engineer controls changes to retrieval or ranking.
Project: review Beacon Supply search relevance
Review the packet
Keep the 47-query denominator visible. The five empty queries need separate investigation: missing index data, tokenization, filters, and permissions can all cause an empty set. For one judged query, three of the first five eligible results are relevant, giving precision at five of 3/5 or 0.6. Seven query-document judgments are disputed and cannot be silently scored as irrelevant. Hold the release comparison until those judgments and the candidate pool are versioned.
Beacon Supply | query set: 47 | catalog snapshot: C-47
42 queries have candidates; 5 have no eligible result
One judged query: 3 relevant in first 5 -> P@5 = 0.6
Seven disputed query-document judgments: review queue
Decision: no aggregate winner from one queryPerformance and review cost
For Q queries with K inspected ranks, constructing a top-K judgment sheet is O(QK) comparisons after retrieval. Human judging dominates elapsed time and grows with uncertain pairs. Reusing the same snapshot and judgment policy makes a comparison possible; changing either one requires a new baseline. A single query-level score is evidence about that query, never a population-wide result.
Common Mistakes
- Do not call an unjudged document irrelevant.
- Do not infer overall search quality from one query.
- Do not publish a ranking win while disputed labels remain in the release slice.
Related lessons
- Search prompts: define query intent and evaluation slices
- Search prompts: judge relevance without treating unknown as wrong
- Search prompts: diagnose candidate coverage and zero results
- Search prompts: calculate ranking metrics on judged pairs
- Search prompts: interpret clicks and gate a ranking release
- Search relevance prompt decisions
- Evaluation sets: measure the failure cases that matter
- Report prompts: reconcile every claim with the packet
