Pair scores do not automatically form safe duplicate groups. A single weak bridge can join two unrelated incident families.
Duplicate clusters: review transitivity, bridges and source identity
Separate pair confidence from cluster identity
If A matches B and B matches C, it does not follow that A and C describe the same incident. B may be a vague message that fits both. Simple connected components will still merge all three. Define a cluster rule: common verified incident key, reviewer approval for cross-group bridges or a representative-medoid check. Keep a reason and model version for every accepted edge. A cluster ID should remain stable across model updates only through an explicit migration.
Find bridge errors early
Inspect edges whose removal splits a large component, clusters with mixed product versions and members with incompatible protected fields. Review the shortest explanation path between a new message and the cluster representative. A broad outage creates thousands of similar tickets, so cap reviewer workload with risk-ranked samples rather than reviewing every pair. Never let one uncertain edge silently swallow a separate privacy or billing incident.
Evaluate at two levels
Use reviewed pair labels for precision and recall, then compare resulting clusters with incident-level truth. Report false merges, false splits, maximum cluster size and changes after retraining. A method can improve pair F1 while making one catastrophic merge. Reserve whole incidents and time windows for audit; copied templates across splits can hide the problem. Equivalence policy defines a valid edge.
Manage revisions and deletions
When a source message is edited or removed, revalidate edges based on that revision and recompute affected clusters. Do not expose a private ticket title through a cluster shared with another tenant. Store access scope on every membership query. The project stages cluster changes for review and keeps a rollback snapshot.
Implementation
def safe_cluster_members(seed_id, accepted_edges, compatible):
members = {seed_id}
frontier = [seed_id]
while frontier:
current = frontier.pop()
for left, right in accepted_edges:
candidate = right if left == current else left if right == current else None
if candidate is None or candidate in members:
continue
if all(compatible(candidate, existing) for existing in members):
members.add(candidate)
frontier.append(candidate)
return members
edges = [("ticket-47", "ticket-82"), ("ticket-82", "ticket-91")]
compatible = lambda candidate, existing: {candidate, existing} != {"ticket-47", "ticket-91"}
assert safe_cluster_members("ticket-47", edges, compatible) == {"ticket-47", "ticket-82"}
Performance and operating cost
This direct scan revisits all e edges for each accepted member and checks compatibility with current members, reaching O(m·e + m²) time for m members; it is a clear audit example, not a large-scale clustering engine. Index adjacency and cache approved constraints for production. Review cost should target high-impact bridge edges, while reporting false merges as a release-blocking metric.
Common Mistakes
- Taking transitive closure of uncertain pair predictions without review.
- Preserving a cluster ID after its source evidence changed.
- Reporting pair F1 while hiding a large wrongful merge.
- Showing private ticket details to another tenant through shared membership.
Read next
- Paraphrase equivalence: make hard negatives change the decision
- Project: stage support-ticket duplicate groups for review
- Text validation: split conversations, duplicates and time together
- Hybrid text ranking with access filters and reranking
- Project: enforce a privacy-safe support-text pipeline
Continue the workflow: Near-duplicate text families: detect overlap without erasing meaning.
