The useful context for a word may sit outside its sentence. Make the evidence window and decision policy reproducible.
Sense decisions: context windows, policy versions and review
Capture the occurrence, not just the word
A ticket title says “hold,” while the message body says “release the authorization after capture.” The intended meaning concerns a payment hold, not a deployment pause. Record the word span, document revision, local sentence, permitted neighboring fields and workflow metadata available at decision time. Do not use a final agent tag assigned later: that would leak the answer into training. Span offsets make each reviewed occurrence recoverable after annotation.
Set a bounded evidence window
A fixed token window is cheap, but it may cut off a definition in the preceding heading or include an unrelated quotation. Compare sentence-only, section-scoped and structured-field contexts on a reviewed set. Keep the winning window policy with the model version. When a document crosses a section boundary, avoid borrowing clues from the next incident. Passage boundaries determine which neighboring text can reasonably support a decision.
Separate candidate ranking from release
Generate plausible senses and rank them against the recorded context. Release a sense only when its score clears the policy threshold and beats the runner-up by the required margin. Otherwise retain candidates and request review; confidence is not evidence if the score has not been calibrated on the current domain. A new inventory version can add candidates and change the margin. Inventory rules govern the NIL state.
Test stability after a vocabulary shift
Re-run a frozen audit after a product launch, renamed queue or documentation rewrite. Break down errors by source field and term, not solely by model average. Keep reviewed examples from before and after the change so a threshold adjustment does not improve one term by harming another. When sense labels drive search, compare retrieval mistakes and access-filter compliance end to end. Human corrections need an audit trail with original prediction, source revision and reviewer policy.
Implementation
def decide_sense(ranked, minimum_score, minimum_margin, inventory_version):
if not ranked:
return {"state": "nil", "inventory": inventory_version}
ordered = sorted(ranked, key=lambda candidate: candidate["score"], reverse=True)
lead = ordered[0]
runner_up = ordered[1]["score"] if len(ordered) > 1 else 0.0
if lead["score"] < minimum_score or lead["score"] - runner_up < minimum_margin:
return {"state": "review", "inventory": inventory_version,
"candidates": [candidate["sense_id"] for candidate in ordered]}
return {"state": "proposed", "inventory": inventory_version,
"sense_id": lead["sense_id"]}
ranked = [{"sense_id": "payment-hold", "score": 0.82},
{"sense_id": "deploy-hold", "score": 0.47}]
assert decide_sense(ranked, 0.75, 0.20, "sense-r4")["sense_id"] == "payment-hold"
Performance and operating cost
Sorting s candidates costs O(s log s) time and O(s) space; a two-best scan would make the decision O(s). Context construction costs O(c) for c inspected tokens. The larger operating cost is maintaining reviewed, domain-specific cases whenever the vocabulary or sense policy changes. A threshold chosen on stale examples can make a cheap classifier expensive in downstream search errors.
Common Mistakes
- Using post-resolution metadata as context for an intake decision.
- Treating an uncalibrated model score as a probability of correctness.
- Letting context cross into an unrelated document section.
- Overwriting the original prediction when a reviewer changes the sense.
