Resolve transliteration candidates with independent record evidence, explicit ties and a reversible review trail.
Cross-script linkage: contextual evidence and collision review
Rank identities, not strings
The same written name can belong to two vendors. A transliteration score measures whether spellings correspond, not whether the business records are identical. Compare permitted candidates using stable evidence: tax registration, verified account number, city, service history or an owner-approved alias. Give each feature a provenance and freshness date. Avoid counting three copied aliases as three independent clues. Variant generation should hand over a candidate set, not an automatic winner.
Define a tie and a NIL path
A decision policy should be able to say “no listed entity” and “two plausible entities.” If the best candidate is only slightly ahead of another, hold for review. If the name looks close but a verified identifier conflicts, reject the link. Choose thresholds on a held-out set containing same-name businesses, script variants and unseen entities. Entity merge review covers corrections when a previously accepted identity is later challenged.
Make decisions reversible
Store the source record, candidate IDs, field evidence, alias model version, reviewer and decision. Keep a merge operation separate from candidate retrieval. If a reviewer corrects a link, update dependent search indexes and derived reports, while preserving the original text and historical decision. Scope each candidate lookup to authorized records; a high score cannot bypass access policy. A confidence value without the compared candidate set is hard to audit.
Test against costly errors
Count false merges separately from missed matches. A missed match creates extra review work; a false merge can combine invoices, permissions or personal data from different organizations. Test recent aliases, short names, punctuation loss and copied addresses. Slice by source script and region because aggregate accuracy can hide a failing subgroup. The vendor intake project requires an explicit review state when evidence is tied.
Implementation
def choose_vendor(candidate_rows, minimum_score=3, margin=2):
ranked = sorted(candidate_rows,
key=lambda row: (-row["evidence_score"], row["entity_id"]))
if not ranked or ranked[0]["evidence_score"] < minimum_score:
return {"state": "nil"}
if len(ranked) > 1 and ranked[0]["evidence_score"] - ranked[1]["evidence_score"] < margin:
return {"state": "review", "candidate_ids":
[row["entity_id"] for row in ranked[:2]]}
return {"state": "linked", "entity_id": ranked[0]["entity_id"]}
vendors = [{"entity_id": "vendor-47", "evidence_score": 5},
{"entity_id": "vendor-82", "evidence_score": 4}]
assert choose_vendor(vendors)["state"] == "review"
assert choose_vendor([{**vendors[1], "evidence_score": 1}, vendors[0]]) == {
"state": "linked", "entity_id": "vendor-47"}
assert choose_vendor([])["state"] == "nil"
Performance and operating cost
Sorting c candidates costs O(c log c) time and O(c) space. The toy score is an input, not a trustworthy model output: production evidence needs calibrated weights, conflict rules, access filtering and reviewer logs before a link changes records.
Common Mistakes
- Calling two copied aliases independent evidence.
- Using a raw similarity score as proof of legal identity.
- Dropping the runner-up candidate from audit logs.
- Applying a merge before a reviewer can inspect conflicting identifiers.
