A similarity score ranks candidate pairs; it is not automatically the probability that two records name the same entity.
Match scores and review bands: separate similarity from identity
Specify fields and failures
For each generated pair, compare source-issued tokens, normalized name, address and verified contact information under explicit data-use rules. A shared address has different evidential value for a single-occupant account than for a household. Missing information is not disagreement. Record which fields were unavailable, which agreed and which conflicted. Candidate generation determines the pairs this stage can see.
Choose bands from reviewed cases
Set a high-confidence acceptance band, a low-confidence rejection band and a middle band for human review. Evaluate false merges and missed links separately because their costs differ: a false merge can expose one customer’s history to another, while a missed link can split one customer’s orders across two rows. The thresholds belong to a versioned policy and should be checked on a held-out reviewed set, including shared-family and name-change cases.
Do not overread the number
A score of 0.83 from a hand-weighted rule is not an 83% chance of a match. Even a calibrated probability depends on the candidate population and review-label quality. If the blocker changes, the mix of candidate pairs changes too, and an old threshold may no longer have the same precision. Calibration is relevant only when a probabilistic score and valid review labels are available.
Protect reviewer independence
Present reviewers with enough evidence to decide, but avoid exposing a proposed score if it will anchor judgment. Record independent decisions, disagreement, adjudication and an uncertain outcome when evidence is insufficient. Do not force ambiguous household records into a binary class. Blinded review can reveal systematic disagreement before a merge policy is released.
Check one awkward pair
Two records share a surname and postal district but have different first names and distinct verified account tokens. A name-similarity feature may still score high, yet the token conflict should block automatic acceptance. The policy needs explicit hard conflicts, not just a weighted total. Keep pair-level reasons so an analyst can explain why the same-looking address did not produce a merge.
Implementation
def review_band(pair_evidence):
if pair_evidence["verified_token_conflict"]:
return "reject"
weighted_fields = ((0.45, pair_evidence["verified_token_agreement"]),
(0.35, pair_evidence["name_similarity"]),
(0.20, pair_evidence["address_similarity"]))
available = [(weight, value) for weight, value in weighted_fields
if value is not None]
if not available:
return "review"
score = sum(weight * value for weight, value in available) / sum(
weight for weight, _ in available)
if score >= 0.91 and pair_evidence["verified_token_agreement"] == 1:
return "accept"
if score <= 0.38:
return "reject"
return "review"
assert review_band({"verified_token_conflict": True,
"verified_token_agreement": 0,
"name_similarity": 0.97,
"address_similarity": 1.0}) == "reject"
assert review_band({"verified_token_conflict": False,
"verified_token_agreement": None,
"name_similarity": 0.97,
"address_similarity": 1.0}) == "review"Performance and operating cost
Scoring P candidate pairs with a fixed number of fields costs O(P) time and O(1) auxiliary space per pair. Producing explanations and review queues requires O(P) storage; reviewer time dominates for pairs in the middle band. This illustrative score is a policy rule, not a calibrated probability.
Common Mistakes
- Do not interpret an arbitrary similarity score as a match probability.
- Do not auto-merge pairs with a documented hard conflict because their soft fields agree.
- Do not evaluate thresholds only on the same reviewed cases used to choose them.
