Build an index that finds inflected support terms while keeping product names and ticket codes exact.
Project: release morphology-aware support search
Set the search task
Agents search a restricted support corpus by issue description, incident ID and exact order code. Keep the original ticket and its access scope as the source of truth. Index surface forms alongside a controlled lemma expansion, not in place of them. Search for “returned item” should reach a relevant “returns” case, while ZX-47 must never match ZX-74. Record the version of the text transform used by each index shard.
Create an adversarial audit
Gather approved queries from agent sessions and write relevance judgments for a fixed top-eight review depth. Include plural forms, past tense, ambiguous words, code-switched notes and product terms that resemble ordinary words. Add exact-code negatives and deleted tickets. Group near-duplicate incidents before splitting so an easy copied template does not inflate the outcome. The term contract guides expansion.
Build and compare
Construct an exact baseline, then a surface-plus-lemma index from the same corpus revision. Compare recall, false expansions, precise code lookup, latency and storage. A contextual tagger can be added only where its incremental gain clears the review cost. Versioned reindexing provides the release and rollback procedure. Recheck tenant filters after candidate generation and before a result is shown.
Ship the guarded index
Keep both index versions during a staged rollout. Inspect misses and unexpected matches from real queries without logging unrestricted ticket text. Exclude deleted sources from both versions. If technical identifiers miss or cross-tenant results appear, return to the old alias immediately. Record source revision, analyzer digest and audit metrics for each promotion.
Implementation
def search_hit_allowed(hit, tenant_id, removed_ids):
return (hit["tenant_id"] == tenant_id
and hit["document_id"] not in removed_ids)
hits = [
{"document_id": "case-47", "tenant_id": "north"},
{"document_id": "case-82", "tenant_id": "south"},
]
visible = [hit for hit in hits if search_hit_allowed(hit, "north", set())]
assert [hit["document_id"] for hit in visible] == ["case-47"]
Performance and operating cost
The access filter is O(1) average per candidate with a set of deleted IDs; scanning k candidates is O(k). The morphology index costs extra postings, and contextual tagging adds per-document and query inference. Compare all of those with the cost of agents missing a relevant case. Access filtering must be enforced independently of model relevance.
Common Mistakes
- Letting a lemma posting override an exact ticket-code match.
- Evaluating with training incidents repeated in the query audit.
- Serving results before tenant and deletion checks.
- Promoting a new analyzer without a matching rebuilt index.
Read next
- Lemmas, part of speech and domain terms in a search index
- Morphology ambiguity, index versions and rollback
- Hybrid text ranking with access filters and reranking
- Project: release an incident-runbook search service
- Project: enforce a privacy-safe support-text pipeline
Continue the workflow: Project: disambiguate support terms before runbook search.
