Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Query rewrites: provenance, protected tokens and user intent

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A search rewrite may improve recall while changing the question. Keep the original query and protect terms that carry operational meaning.

Distinguish correction from expansion

Correcting a typo in “authorizaton” differs from adding “payment” to “hold.” The first may recover an intended token; the second chooses a domain. Record the original query, proposed rewrite, rule or model version, affected spans and reason. Never replace a quoted phrase, incident ID, error code or negated term without explicit policy. Domain sense decisions can support an expansion, but an uncertain sense should not become an invisible hard filter.

Use a bounded candidate set

Generate a few alternatives from approved terminology, spelling rules or reviewed query patterns. Keep the unchanged query in retrieval, and treat additions as optional recall paths until judged cases show they help. An expansion from a broad synonym list can pull in an unrelated domain and displace the exact match. Scope candidates by language and product, and record why each term was added. Hybrid ranking can combine original and expanded matches while still enforcing document access.

Preserve operators and exclusions

The query “refund not pending” should not become a search for pending refunds. Preserve negation and quoted spans in a structured representation. A minus operator, field filter or date constraint may be syntax rather than prose. Parse the application’s supported query grammar before rewriting its free-text portion. Unknown syntax should pass through unchanged or request user clarification, not be guessed. A rewrite is a proposal about the user’s intent, not a fact about the corpus.

Measure the whole result page

Compare original-only and rewritten results on incident-scoped judged queries. Report correct-runbook recall, wrong-domain top results, latency and the percentage of queries changed. Slice by short queries, identifiers, quoted text and negatives. Show reviewers both query forms and the matched evidence. A rewrite that increases recall but hides the only exact answer is a regression. Feedback audits help identify misleading post-release signals.

Implementation

python
APPROVED_ALIASES = {"authorization": ("card authorization",),
                    "rollback": ("revert deployment",)}

def rewrite_query(query, protected_tokens):
    terms = query.split()
    optional = []
    for term in terms:
        if term in protected_tokens or term.startswith("-") or term.isdigit():
            continue
        optional.extend(APPROVED_ALIASES.get(term.lower(), ()))
    return {"original": query, "optional_phrases": tuple(optional)}

proposal = rewrite_query("authorization -pending INC47", {"INC47"})
assert proposal["original"] == "authorization -pending INC47"
assert proposal["optional_phrases"] == ("card authorization",)

Performance and operating cost

Scanning t space-separated terms is O(t) time, excluding the size of emitted expansions. A real search grammar needs a proper parser for quotes, escapes and filters; this code deliberately emits optional phrases without altering the original. Every added branch can increase retrieval and ranking work, so cap candidate count and measure tail latency with relevance.

Common Mistakes

  • Overwriting the original query with an unreviewed guess.
  • Expanding inside an incident ID or quoted phrase.
  • Dropping negation while normalizing text.
  • Judging success from recall alone when wrong-domain top results rise.

Read next

ai-data
natural-language-processing
Storage details