Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: disambiguate support terms before runbook search

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Stage sense decisions for ambiguous support vocabulary, then test whether the correct runbook remains visible without false joins.

Define the release boundary

A support search receives “charge failed after swap” and must distinguish a failed card debit from battery charging. It may rank documents, but it must not execute refunds, change a device or bypass access filters. Create a small inventory for charge, hold and retry with stable IDs, exclusions and NIL. Reviewers annotate exact spans and document revisions. The sense inventory provides the label contract.

Build difficult evaluation cases

Collect tickets and runbook excerpts from billing, hardware and deployment queues. Include short titles, copied text, quotations, negated statements and cases where the decisive clue is in a section heading. Split by incident and copied template so near-duplicates do not inflate quality. Freeze the inventory version and reviewer guidelines. Keep an independent holdout with examples of new meanings and ambiguous cases that should abstain rather than force a label.

Stage the search decision

Produce candidates with the text span and allowed context. Apply a margin policy, retaining NIL and review states. Pass a proposed sense to ranking as an auditable feature, while enforcing document ACLs before results appear. Compare sense-aware ranking with the existing search baseline; never hide the only correct runbook solely because the classifier was uncertain. Context and version policy fixes what each model saw.

Measure the user-visible error

Report top-result relevance, wrong-domain results, correct-runbook recall, NIL rate, reviewer time and errors on rare senses. Slice by term, queue and document age. Shadow the change before rollout, then sample failures after vocabulary updates. Retain source, prediction and reviewer correction so a new sense definition can be tested against old cases. Roll back the ranking feature if wrong-domain results rise, even when sense accuracy appears stable.

Implementation

python
def rank_runbooks(query_sense, candidates, allowed_document_ids):
    visible = [candidate for candidate in candidates
               if candidate["document_id"] in allowed_document_ids]
    ranked = sorted(visible,
                    key=lambda candidate: (
                        candidate["base_score"] +
                        (0.17 if query_sense == candidate["sense_id"] else 0.0)),
                    reverse=True)
    return [candidate["document_id"] for candidate in ranked]

runbooks = [{"document_id": "billing-47", "base_score": 0.61,
             "sense_id": "payment-debit"},
            {"document_id": "device-82", "base_score": 0.69,
             "sense_id": "battery-energy"}]
assert rank_runbooks("payment-debit", runbooks, {"billing-47", "device-82"})[0] == "billing-47"
assert rank_runbooks("payment-debit", runbooks, {"device-82"}) == ["device-82"]

Performance and operating cost

Filtering n candidates and sorting the visible set costs O(n log n) time and O(n) space. Sense detection and reviewer handling add separate costs. The fixed boost is illustrative and needs validation against the frozen holdout; it is not a substitute for access control. Track whether uncertain queries lose the right result rather than celebrating a higher average ranking score.

Common Mistakes

  • Applying a sense boost before enforcing document access.
  • Suppressing all results when a word has no confident sense.
  • Using tickets from one incident on both sides of evaluation.
  • Treating improved label accuracy as proof that search improved.

Read next

ai-data
natural-language-processing
Storage details