Clicks reflect exposure as well as relevance. Evaluate reformulations against independent judgments before turning interaction traces into policy.
Search feedback: position bias, query drift and audit design
Separate observation from judgment
A user may click the first result because it is visible, not because it answers the question. A zero-click session may mean the snippet answered it or the results failed. Log query form, displayed positions, eligible document set and action timing; do not equate a click with a relevant label. Protect user and document identifiers under retention policy. Rewrite provenance identifies which query version produced the result page.
Watch for feedback loops
If an expanded query promotes one runbook, that runbook receives more exposure and therefore more clicks. Training the next expansion on those clicks can reinforce its own initial error. Keep a frozen set of independently judged queries and compare it across releases. Randomized exploration can estimate exposure effects where permitted, but it still needs privacy and safety review. An operator’s explicit relevance mark is stronger evidence than a bare click, though it can still reflect task context.
Detect movement away from intent
A query for “card hold release” may expand toward generic refund content because the corpus contains many refund articles. Track whether newly introduced terms dominate top results and whether exact-match evidence disappears. Compare original-only, expanded-only and combined rankings. Flag a rewrite when it moves the correct runbook below the first screen or brings in a wrong-domain answer. Abbreviation collisions can make this drift worse for short queries.
Build a release audit
Group related incident queries, copied templates and follow-up reformulations together before splitting. Ask reviewers to judge result relevance without seeing which algorithm produced it. Measure top-result errors, correct-result recall, query-change rate and latency by language and product. After release, sample changed queries rather than only popular ones. Roll back the rewrite rule when the audited error rises, even if overall click-through increases.
Implementation
def audit_changed_results(original_rank, rewritten_rank, relevant_ids):
if not original_rank or not rewritten_rank:
return {"review": True, "reason": "empty-result-list"}
original_hit = original_rank[0] in relevant_ids
rewritten_hit = rewritten_rank[0] in relevant_ids
return {"review": original_hit and not rewritten_hit,
"original_top_relevant": original_hit,
"rewritten_top_relevant": rewritten_hit}
audit = audit_changed_results(["runbook-47", "note-82"],
["note-82", "runbook-47"], {"runbook-47"})
assert audit["review"] and not audit["rewritten_top_relevant"]
Performance and operating cost
The top-result check is O(1) with a set of judged relevant IDs, but collecting unbiased judgments is the expensive part. Auditing all k ranked results costs O(k) per query and exposes more regressions. Keep the judged set independent of click-derived training signals; otherwise a feedback loop can make a weak policy appear to improve itself.
Common Mistakes
- Using raw clicks as ground-truth relevance labels.
- Comparing rankings without preserving exposure positions.
- Allowing rewritten terms to displace the only exact answer.
- Splitting follow-up queries from one incident across train and test.
