A search evaluation slice is a fixed set of queries grouped by the task the user meant to complete. Record normalized query text, language, device or role when relevant, intent class, catalog snapshot, filters, and expected eligible result type. A prompt may propose intent labels, but it must expose ambiguous queries and preserve the raw wording. A slice drawn only from successful searches cannot measure empty searches or uncommon terminology. The denominator belongs to the query set, not the number of results retrieved.
Search prompts: define query intent and evaluation slices
Operational case
Beacon Supply freezes 47 help and catalog queries at snapshot C-47. Five have no eligible result; they stay in the slice. A query for 'seal 47' may mean a part number or an installation guide, so reviewers keep both candidate intents until a session or owner clarifies the task. The model writes a proposed label and its evidence, not a made-up purchase intent. Role-specific filters are logged with the query because they change what could have been returned.
query_id: BS-Q47
raw_text: seal 47
intent_candidates: part lookup | installation help
catalog_snapshot: C-47
role_filter: contractor
intent_status: reviewPerformance and review cost
Grouping Q queries by a small intent taxonomy is O(Q) record work plus human review for ambiguous cases. Snapshotting the eligible catalog costs storage proportional to indexed records, but prevents later catalog changes from rewriting the test. Retain rare and zero-result queries even if their count makes a headline score less flattering.
Common Mistakes
- Do not discard empty queries from the denominator.
- Do not label ambiguity as certainty because one result looks plausible.
- Do not compare slices built from different catalog snapshots.
Connected lessons
- Prompt engineering applications
- Prompt Engineering
- Evaluation sets: measure the failure cases that matter
- Retrieved context: select sufficient evidence before writing the answer
- Search prompts: judge relevance without treating unknown as wrong
- Search prompts: diagnose candidate coverage and zero results
- Search prompts: calculate ranking metrics on judged pairs
- Search prompts: interpret clicks and gate a ranking release
- Project: review Beacon Supply search relevance
- Search relevance prompt decisions
