A click records exposure and user action, not a clean relevance grade. Position, presentation, prior familiarity, and the absence of a result all affect whether a user clicks. A release review combines a fixed judged set with live signals, segment breakdowns, zero-result rate, permissions checks, and latency. The prompt should draft a decision memo from named measurements and state which evidence is missing. It cannot turn an attractive click-through change into authorization to publish or ignore a safety gate.
Search prompts: interpret clicks and gate a ranking release
Operational case
Beacon's proposed ranker improves clicks on popular part searches but has no adjudicated result for seven disputed pairs, while five empty queries remain under investigation. A mobile contractor segment also has a slower result response. The release memo records the observed click change as a hypothesis, lists the unresolved judgments, and requests segment and latency checks. The engineer holds the rollout until the agreed gates pass; the model neither deploys nor calls clicks proof of relevance.
offline: 7 disputed judgments pending
empty queries: 5 / 47, causes tracked separately
online: clicks by segment, exposure, and position
guards: permission, no-result rate, latency
release decision: hold until named gates passPerformance and review cost
Aggregating E exposure events is O(E) scan work, with additional storage for segment counters. Judged-set comparisons cost O(QK) as configured. Recording exposure and position adds telemetry volume but prevents a misleading interpretation of clicks. Release gates should use measured thresholds chosen before reading the outcome.
Common Mistakes
- Do not use clicks as unbiased relevance labels.
- Do not average away a failing user segment.
- Do not invent a release threshold after seeing the result.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Online prompt experiments: define exposure and stop rules first
- Analytics prompts: define the metric and event contract
- Search prompts: define query intent and evaluation slices
- Search prompts: judge relevance without treating unknown as wrong
- Search prompts: diagnose candidate coverage and zero results
- Search prompts: calculate ranking metrics on judged pairs
- Project: review Beacon Supply search relevance
- Search relevance prompt decisions
