Build a scoped intake review that retains source-script names, generates candidates and refuses a merge when independent evidence is weak.
Project: link vendor names across scripts without false merges
Assemble the fixture
Prepare vendor records with Devanagari and Latin name variants, two businesses that share a short trading name, an unseen business and one stale address. Include a tenant boundary. Each incoming mention needs a source text offset, document revision, script label and permitted candidate set. Use fictional identifiers and consistent field values. Variant generation supplies recall without pretending that the first spelling is authoritative.
Write the gold decision
For every intake record, annotate linked, NIL or review. A reviewer should explain which independent field resolved a tie: a verified registration number can support a link, while a copied address alone cannot. Freeze the annotation policy and keep corrections by version. Split evaluation by underlying vendor rather than name string so a variant of one vendor cannot leak across train and test. Measure reviewer disagreement as a policy signal, not just model noise.
Implement an admission boundary
Lookup may produce several candidates, but the project must only admit candidates inside the authenticated tenant. A matching alias outside that set remains inaccessible. Reject a proposed merge when a verified identifier conflicts, even if the transliterated name looks perfect. The code below represents the final admission check; collision review supplies the ranking and tie state before this check.
Release by error cost
Report candidate recall, linked precision, false merges, NIL accuracy, review volume and median reviewer time. Inspect each false merge with its source text and evidence trail. Re-run the fixture after alias updates and test that earlier links remain stable or are explicitly invalidated. Do not silently overwrite billing or access records during a name reconciliation; a project can ship candidate suggestions before it ships automated merges.
Implementation
def admit_vendor_link(decision, registration, tenant_vendor_ids):
if decision.get("state") != "linked":
return {"state": "hold", "reason": "unresolved-link"}
vendor_id = decision["entity_id"]
if vendor_id not in tenant_vendor_ids:
return {"state": "deny", "reason": "tenant-scope"}
expected = registration.get(vendor_id)
if expected is None or expected != decision.get("registration_id"):
return {"state": "hold", "reason": "identifier-conflict"}
return {"state": "admitted", "entity_id": vendor_id}
registration = {"vendor-47": "reg-47", "vendor-82": "reg-82"}
proposal = {"state": "linked", "entity_id": "vendor-47",
"registration_id": "reg-47"}
assert admit_vendor_link(proposal, registration, {"vendor-47"})["state"] == "admitted"
assert admit_vendor_link({**proposal, "registration_id": "reg-82"},
registration, {"vendor-47"})["state"] == "hold"
Performance and operating cost
With hash-backed sets and maps, this admission check is expected O(1) time and space per proposal. Candidate generation and evidence review cost more and depend on catalog size. The snippet verifies a supplied identifier; production must authenticate its source and ensure the registry is current before changing any record.
Common Mistakes
- Letting a name match merge records across tenants.
- Using one gold spelling as the only accepted rendering.
- Treating an unresolved tie as a low-confidence link instead of review.
- Mutating billing records before measuring false merges.
