Two support messages can share almost every word while disagreeing about status, amount or actor. Label equivalence for the downstream decision, not surface overlap.
Paraphrase equivalence: make hard negatives change the decision
Write the equivalence rule
For duplicate incident detection, two messages are equivalent only when they describe the same operational issue at a compatible time and scope. “The gateway retried 47 times” and “The gateway allows 47 retries” are related but not duplicates: observed count and configured limit answer different questions. Store source revisions, conversation group, entity IDs and a reviewer label with an explicit reason. A semantic similarity score is a candidate signal, not a truth label. Entailment asks whether one statement supports another; bidirectional equivalence is stricter.
Build adversarial pairs
Include changed negation, numbers, product versions, speakers and dates. Add truly equivalent pairs with different wording, word order and abbreviations. A random unrelated negative set is too easy. Review mined hard negatives because an apparently different message may be a legitimate duplicate after context is read. Keep copied templates and repeated incidents in one split. Grouped temporal validation prevents the same event from leaking into train and audit.
Handle asymmetry and context
“Payment failed” may summarize “payment failed after the callback timed out,” but the shorter sentence omits a cause; whether that counts as a duplicate depends on the task. Define whether paraphrase requires exact informational equivalence or merely the same resolution queue. Do not silently change between these definitions. Record surrounding context needed to resolve pronouns and ticket references, while withholding private text from users without access.
Measure the actual decision
Report pair precision and recall at a chosen threshold, then evaluate cluster errors once pairs are connected. A few false pair links can merge unrelated incidents. Break out number changes, negation and cross-language pairs. Compare a lexical baseline against an embedding candidate system and a pairwise reranker. Cluster review addresses the nontransitive edges, and the support duplicate project tests the operational release.
Implementation
def protected_fields_match(left, right):
if left["incident_id"] != right["incident_id"]:
return False
for field in ("observed_status", "product_version", "amount_minor"):
if left.get(field) != right.get(field):
return False
return True
first = {"incident_id": "inc-47", "observed_status": "failed",
"product_version": "v7", "amount_minor": 2847}
changed = {**first, "amount_minor": 2947}
assert not protected_fields_match(first, changed)
Performance and operating cost
A protected-field gate is O(1) for a fixed field list. Comparing every pair of n messages is O(n²), so generate candidates by time, tenant and index before applying a costly pair model. Candidate pruning reduces work but can lose true duplicates. Publish candidate recall, pair precision and wrongful cluster merges alongside p95 matching latency.
Common Mistakes
- Treating shared keywords as equivalence.
- Using an easy random-negative test set and missing changed-number failures.
- Equating one-way entailment with a paraphrase.
- Scoring pair labels without measuring the incident clusters users see.
Read next
- Duplicate clusters: review transitivity, bridges and source identity
- Project: stage support-ticket duplicate groups for review
- Textual entailment: premise, hypothesis and evidence boundaries
- Text embeddings: pair labels, hard negatives and versioned vectors
- Text validation: split conversations, duplicates and time together
