Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: release morphology-aware support search

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Build an index that finds inflected support terms while keeping product names and ticket codes exact.

Set the search task

Agents search a restricted support corpus by issue description, incident ID and exact order code. Keep the original ticket and its access scope as the source of truth. Index surface forms alongside a controlled lemma expansion, not in place of them. Search for “returned item” should reach a relevant “returns” case, while ZX-47 must never match ZX-74. Record the version of the text transform used by each index shard.

Create an adversarial audit

Gather approved queries from agent sessions and write relevance judgments for a fixed top-eight review depth. Include plural forms, past tense, ambiguous words, code-switched notes and product terms that resemble ordinary words. Add exact-code negatives and deleted tickets. Group near-duplicate incidents before splitting so an easy copied template does not inflate the outcome. The term contract guides expansion.

Build and compare

Construct an exact baseline, then a surface-plus-lemma index from the same corpus revision. Compare recall, false expansions, precise code lookup, latency and storage. A contextual tagger can be added only where its incremental gain clears the review cost. Versioned reindexing provides the release and rollback procedure. Recheck tenant filters after candidate generation and before a result is shown.

Ship the guarded index

Keep both index versions during a staged rollout. Inspect misses and unexpected matches from real queries without logging unrestricted ticket text. Exclude deleted sources from both versions. If technical identifiers miss or cross-tenant results appear, return to the old alias immediately. Record source revision, analyzer digest and audit metrics for each promotion.

Implementation

python
def search_hit_allowed(hit, tenant_id, removed_ids):
    return (hit["tenant_id"] == tenant_id
            and hit["document_id"] not in removed_ids)

hits = [
    {"document_id": "case-47", "tenant_id": "north"},
    {"document_id": "case-82", "tenant_id": "south"},
]
visible = [hit for hit in hits if search_hit_allowed(hit, "north", set())]
assert [hit["document_id"] for hit in visible] == ["case-47"]

Performance and operating cost

The access filter is O(1) average per candidate with a set of deleted IDs; scanning k candidates is O(k). The morphology index costs extra postings, and contextual tagging adds per-document and query inference. Compare all of those with the cost of agents missing a relevant case. Access filtering must be enforced independently of model relevance.

Common Mistakes

  • Letting a lemma posting override an exact ticket-code match.
  • Evaluating with training incidents repeated in the query audit.
  • Serving results before tenant and deletion checks.
  • Promoting a new analyzer without a matching rebuilt index.

Read next

Continue the workflow: Project: disambiguate support terms before runbook search.

ai-data
natural-language-processing
Storage details