Build a duplicate-review queue that proposes incident groups, blocks incompatible status or amount changes and keeps every merge reversible.
Project: stage support-ticket duplicate groups for review
Set the workflow
Support agents want to see potentially repeated customer reports without losing separate cases. The service proposes a candidate pair with reasons and protected-field comparisons; a reviewer accepts a merge or rejects it. No model output deletes a ticket or collapses customer histories automatically. Keep tenant scope and incident permissions attached to all candidates.
Construct a hard corpus
Collect messages from the same incident plus near misses from adjacent incidents. Include number changes, negated status, repeated templates, revised ticket text and messages from different customer accounts. Label pair equivalence under one policy version. Group by incident and time before splitting. Use the equivalence contract to keep an identical phrase from overriding an incompatible observed fact.
Stage cluster changes
Retrieve likely pairs, run a pair model and protected-field gate, then propose an edge. Compute the candidate cluster and flag any bridge that would unite two previously separate groups. Require a reviewer for high-impact merges. Cluster review explains why pair accuracy alone is insufficient. Store before and after membership snapshots so a mistaken merge can be undone.
Operate safely
Track reviewer acceptance, false merges, false splits, queue age and cross-tenant access checks. Sample confident automatic suggestions for audit even if they were not merged. A new embedding model needs a fresh index and replay against the fixed incident-level audit set. If a ticket is deleted under retention policy, remove derived vectors and invalidate its incident-group edges.
Implementation
def stage_duplicate_edge(left, right, same_tenant, fields_compatible):
if left == right:
raise ValueError("a ticket cannot duplicate itself")
if not same_tenant:
return {"state": "rejected", "reason": "tenant-boundary"}
if not fields_compatible:
return {"state": "review", "reason": "protected-field-conflict"}
return {"state": "review", "reason": "candidate-duplicate",
"tickets": tuple(sorted((left, right)))}
proposal = stage_duplicate_edge("ticket-47", "ticket-82", True, True)
assert proposal["state"] == "review"
Performance and operating cost
The staging gate is O(1) time and space. Candidate retrieval and pair reranking dominate compute; incident-level review dominates human cost. Model throughput is secondary to wrongful merge rate. Store only the minimum pair evidence required for review, and enforce tenant access before the candidate is displayed or cached.
Common Mistakes
- Auto-merging tickets because their pair score exceeds a threshold.
- Ignoring a changed amount or negated status.
- Testing pairs while never measuring group-level false merges.
- Leaving deleted ticket embeddings in the candidate index.
Read next
- Paraphrase equivalence: make hard negatives change the decision
- Duplicate clusters: review transitivity, bridges and source identity
- Text embeddings: pair labels, hard negatives and versioned vectors
- Project: enforce a privacy-safe support-text pipeline
- Text validation: split conversations, duplicates and time together
