A merged entity graph changes denominators, histories and permissions; every released version needs a traceable decision path.
Entity merge lineage: show how identity decisions change metrics
Review clusters, not only edges
If A matches B and B matches C, a transitive grouping step places A, B and C together even if A and C have a strong conflict. Pair acceptance is therefore not the last quality gate. Inspect component size, contradictory tokens and cross-tenant membership before publishing a cluster. Pair review supplies approved edges, but a cluster policy must decide whether their transitive closure is valid.
Version stable entity IDs
An entity ID should identify the released cluster under a declared resolution version. Keep source record IDs and a mapping from each to its entity ID. When a false merge is split, retain old-to-new lineage rather than overwriting all history. A metric computed under release 4 may count 46 customers while release 5 counts 47 from the same raw orders. Both results can be reproducible if the identity release is pinned.
Recompute downstream rates
Repeat-purchase rate uses distinct customers as its denominator. A false merge can create an apparent repeat buyer by combining two people’s first orders; a missed link can hide a repeat buyer as two one-time buyers. After any material merge revision, recompute affected counts and publish the difference by reason. Retention cohorts may also shift because the first-purchase date for a merged entity changes.
Protect access boundaries
Entity resolution can join records that previously lived in separate privacy or tenant scopes. A technically plausible match is not authorization to reveal combined histories. Preserve purpose, consent and access checks at the resolved-entity layer, then test a false-merge rollback. Data minimization remains relevant even when the use case is analysis rather than model training.
Release with a diff
For each new identity version, report linked and unlinked record counts, cluster-size distribution, accepted-edge reasons, manual review volume, split and merge counts, and downstream metric restatements. Quarantine giant or contradictory components. A reviewer should trace one customer count from source rows through pair decisions to the published entity mapping.
Implementation
def entity_components(record_ids, approved_links):
parent = {record_id: record_id for record_id in record_ids}
def find(record_id):
while parent[record_id] != record_id:
parent[record_id] = parent[parent[record_id]]
record_id = parent[record_id]
return record_id
for left_id, right_id in approved_links:
if left_id not in parent or right_id not in parent:
raise ValueError("approved link references an unknown record")
parent[find(right_id)] = find(left_id)
components = {}
for record_id in record_ids:
components.setdefault(find(record_id), set()).add(record_id)
return list(components.values())
assert entity_components(["ret-7", "sup-7", "ret-9"],
[("ret-7", "sup-7")]) == [{"ret-7", "sup-7"}, {"ret-9"}]Performance and operating cost
With path compression, building components from N records and E approved links is near O(N + E) in practice and uses O(N) memory. Cluster-level conflict checks add comparisons within components; they are necessary before a transitive merge is trusted.
Common Mistakes
- Do not assume two approved pair links imply a safe three-record cluster.
- Do not overwrite old entity mappings without a release lineage.
- Do not let a plausible match bypass tenant or consent boundaries.
