Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Paraphrase equivalence: make hard negatives change the decision

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Two support messages can share almost every word while disagreeing about status, amount or actor. Label equivalence for the downstream decision, not surface overlap.

Write the equivalence rule

For duplicate incident detection, two messages are equivalent only when they describe the same operational issue at a compatible time and scope. “The gateway retried 47 times” and “The gateway allows 47 retries” are related but not duplicates: observed count and configured limit answer different questions. Store source revisions, conversation group, entity IDs and a reviewer label with an explicit reason. A semantic similarity score is a candidate signal, not a truth label. Entailment asks whether one statement supports another; bidirectional equivalence is stricter.

Build adversarial pairs

Include changed negation, numbers, product versions, speakers and dates. Add truly equivalent pairs with different wording, word order and abbreviations. A random unrelated negative set is too easy. Review mined hard negatives because an apparently different message may be a legitimate duplicate after context is read. Keep copied templates and repeated incidents in one split. Grouped temporal validation prevents the same event from leaking into train and audit.

Handle asymmetry and context

“Payment failed” may summarize “payment failed after the callback timed out,” but the shorter sentence omits a cause; whether that counts as a duplicate depends on the task. Define whether paraphrase requires exact informational equivalence or merely the same resolution queue. Do not silently change between these definitions. Record surrounding context needed to resolve pronouns and ticket references, while withholding private text from users without access.

Measure the actual decision

Report pair precision and recall at a chosen threshold, then evaluate cluster errors once pairs are connected. A few false pair links can merge unrelated incidents. Break out number changes, negation and cross-language pairs. Compare a lexical baseline against an embedding candidate system and a pairwise reranker. Cluster review addresses the nontransitive edges, and the support duplicate project tests the operational release.

Implementation

python
def protected_fields_match(left, right):
    if left["incident_id"] != right["incident_id"]:
        return False
    for field in ("observed_status", "product_version", "amount_minor"):
        if left.get(field) != right.get(field):
            return False
    return True

first = {"incident_id": "inc-47", "observed_status": "failed",
         "product_version": "v7", "amount_minor": 2847}
changed = {**first, "amount_minor": 2947}
assert not protected_fields_match(first, changed)

Performance and operating cost

A protected-field gate is O(1) for a fixed field list. Comparing every pair of n messages is O(n²), so generate candidates by time, tenant and index before applying a costly pair model. Candidate pruning reduces work but can lose true duplicates. Publish candidate recall, pair precision and wrongful cluster merges alongside p95 matching latency.

Common Mistakes

  • Treating shared keywords as equivalence.
  • Using an easy random-negative test set and missing changed-number failures.
  • Equating one-way entailment with a paraphrase.
  • Scoring pair labels without measuring the incident clusters users see.

Read next

ai-data
natural-language-processing
Storage details