Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Duplicate clusters: review transitivity, bridges and source identity

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Pair scores do not automatically form safe duplicate groups. A single weak bridge can join two unrelated incident families.

Separate pair confidence from cluster identity

If A matches B and B matches C, it does not follow that A and C describe the same incident. B may be a vague message that fits both. Simple connected components will still merge all three. Define a cluster rule: common verified incident key, reviewer approval for cross-group bridges or a representative-medoid check. Keep a reason and model version for every accepted edge. A cluster ID should remain stable across model updates only through an explicit migration.

Find bridge errors early

Inspect edges whose removal splits a large component, clusters with mixed product versions and members with incompatible protected fields. Review the shortest explanation path between a new message and the cluster representative. A broad outage creates thousands of similar tickets, so cap reviewer workload with risk-ranked samples rather than reviewing every pair. Never let one uncertain edge silently swallow a separate privacy or billing incident.

Evaluate at two levels

Use reviewed pair labels for precision and recall, then compare resulting clusters with incident-level truth. Report false merges, false splits, maximum cluster size and changes after retraining. A method can improve pair F1 while making one catastrophic merge. Reserve whole incidents and time windows for audit; copied templates across splits can hide the problem. Equivalence policy defines a valid edge.

Manage revisions and deletions

When a source message is edited or removed, revalidate edges based on that revision and recompute affected clusters. Do not expose a private ticket title through a cluster shared with another tenant. Store access scope on every membership query. The project stages cluster changes for review and keeps a rollback snapshot.

Implementation

python
def safe_cluster_members(seed_id, accepted_edges, compatible):
    members = {seed_id}
    frontier = [seed_id]
    while frontier:
        current = frontier.pop()
        for left, right in accepted_edges:
            candidate = right if left == current else left if right == current else None
            if candidate is None or candidate in members:
                continue
            if all(compatible(candidate, existing) for existing in members):
                members.add(candidate)
                frontier.append(candidate)
    return members

edges = [("ticket-47", "ticket-82"), ("ticket-82", "ticket-91")]
compatible = lambda candidate, existing: {candidate, existing} != {"ticket-47", "ticket-91"}
assert safe_cluster_members("ticket-47", edges, compatible) == {"ticket-47", "ticket-82"}

Performance and operating cost

This direct scan revisits all e edges for each accepted member and checks compatibility with current members, reaching O(m·e + m²) time for m members; it is a clear audit example, not a large-scale clustering engine. Index adjacency and cache approved constraints for production. Review cost should target high-impact bridge edges, while reporting false merges as a release-blocking metric.

Common Mistakes

  • Taking transitive closure of uncertain pair predictions without review.
  • Preserving a cluster ID after its source evidence changed.
  • Reporting pair F1 while hiding a large wrongful merge.
  • Showing private ticket details to another tenant through shared membership.

Read next

Continue the workflow: Near-duplicate text families: detect overlap without erasing meaning.

ai-data
natural-language-processing
Storage details