Propose bounded search alternatives, preserve the operator’s exact query and reject rewrites that bury the correct runbook.
Project: release controlled runbook query reformulation
Define the search contract
An on-call engineer enters “authorization timeout INC47.” The search service may suggest a reviewed phrase such as “card authorization,” but must preserve INC47, the original terms and any explicit filters. Access rules apply to every result, regardless of rewrite. Store the original query and all optional expansions with a policy version so a result can be explained later. The rewrite contract fixes protected tokens and the audit record.
Create a hard evaluation set
Include terse error-code queries, misspellings, abbreviations with multiple meanings, quoted phrases, negated searches and incidents with several similarly named services. Group all reformulations from one incident in the same split. Reviewers judge which runbook answers each query, including cases with no eligible document. Freeze the judged set before choosing expansion rules, otherwise the team can tune to a remembered list of successes.
Stage results safely
Retrieve using both the original and approved optional phrases, enforce access, then rank with explicit preference for exact evidence. If an expansion changes the top result, record the reason for audit. When all candidates are uncertain, show original-only results and a suggestion the engineer may choose. Do not use clicks from the same staged release as ground-truth training labels. Feedback auditing detects position bias and query drift.
Gate and roll back the release
Compare top-result correctness, correct-runbook recall, wrong-domain rate, access violations and tail latency against original-only search. Inspect every case where a previously correct top result is displaced. Shadow the proposed policy before enabling it, then sample changed queries after release by product and language. Retain the previous alias table and ranking version for rollback. A rise in clicks cannot overrule independently judged wrong answers.
Implementation
def choose_visible_results(original_hits, expanded_hits, allowed_ids):
seen = set()
visible = []
for document_id in original_hits + expanded_hits:
if document_id in allowed_ids and document_id not in seen:
seen.add(document_id)
visible.append(document_id)
return visible
result = choose_visible_results(["runbook-47", "restricted-91"],
["runbook-47", "note-82"],
{"runbook-47", "note-82"})
assert result == ["runbook-47", "note-82"]
Performance and operating cost
Merging n original and m expanded hits is O(n + m) expected time and space with hash sets. Retrieval and reranking add their own latency, so cap expansions and audit the slow tail, not only the median. The order in this illustration favors original hits; a production ranker still needs judged-query evaluation and the same access gate.
Common Mistakes
- Letting expansion results bypass access filtering.
- Replacing the original query instead of retaining it.
- Training on clicks produced by the same ranking policy.
- Calling the release successful while previously correct top answers disappear.
