Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: resolve customer records and audit the metric restatement

Last updated: 5 Oct 20265 min read
project
IntermediateBy AITrove Editorial

Build a versioned customer mapping from two systems, with candidate recall, review decisions, cluster checks and a repeat-purchase diff.

Construct two imperfect sources

Create 47 retail account records and 39 support records. Include a moved customer, a changed family name, two members of one household, reused local IDs in different tenants, one missing postal district and a repeated event ID. Provide a small adjudicated set of known matches and known nonmatches, including difficult cases. Raw records must remain available. Namespace rules determine exact matches.

Generate candidates in passes

Resolve exact issued tokens first. Then produce candidate pairs from two independent blocking keys and deduplicate pair IDs. Report pair count, maximum block size and recall against the reviewed true-match set. Include a true match missed by one pass but caught by the other. A known match missed by both is a blocker failure, not a scoring failure. The blocking lesson sets that audit.

Review and release clusters

Run a versioned score rule with hard conflicts. Put ambiguous pairs in a reviewer queue and keep review evidence. Build components only from approved edges, then fail components with contradictory verified tokens or mixed tenants. Give each released component a stable versioned entity ID. Merge lineage must support a later split.

Show the analytical consequence

Compute distinct buyers and repeat-purchase rate from the same frozen order table under identity release A and release B. Report added, removed and changed entity assignments and the exact numerator and denominator difference. A false household merge should create a detectable artificial repeat buyer; correcting it should restate the metric with an explanation, not silently change a dashboard.

Deliver proof of reproducibility

Save raw fixtures, canonicalization rules, candidate pairs, reviewed labels, approved edges, cluster conflicts, released mappings, metric outputs, code revision and snapshot cutoff. Add assertions for tenant isolation, candidate recall on the reviewed set and exact reconciliation of all source record IDs into either a released entity or a documented unresolved state.

Implementation

python
def repeat_buyer_count(order_rows, entity_by_account):
    orders_by_entity = {}
    for order in order_rows:
        entity_id = entity_by_account[order["account_id"]]
        orders_by_entity.setdefault(entity_id, set()).add(order["order_id"])
    return sum(len(order_ids) >= 2 for order_ids in orders_by_entity.values())

orders = [{"account_id": "ret-17", "order_id": "ord-41"},
          {"account_id": "ret-18", "order_id": "ord-42"}]
assert repeat_buyer_count(orders, {"ret-17": "person-A", "ret-18": "person-B"}) == 0
assert repeat_buyer_count(orders, {"ret-17": "person-A", "ret-18": "person-A"}) == 1

Performance and operating cost

The final metric scan costs O(O) expected time and space for O orders. Candidate generation can still create quadratic work within large blocks; pair review, cluster validation and privacy checks dominate the operating cost. Preserve an identity-release key with every published metric.

Common Mistakes

  • Do not score pairs that blocking never generated and call them rejected.
  • Do not publish a transitive cluster without checking contradictions.
  • Do not compare repeat rates across identity releases without a restatement ledger.

Read next

ai-data
data-science
Storage details