Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Transliteration variants are candidates, not identity

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Treat names written in another script as versioned candidate evidence; a similar rendering does not establish that two records describe one entity.

Separate the written form from the entity

Transliteration maps a name between writing systems. It does not translate an organization, prove a legal identity or guarantee a unique spelling. A vendor registered as “सारिका” might appear as “Sarika” in one import and “Saarika” in another. Keep the source spelling, script, language tag, span offsets and record identifier. Do not replace the original with a Latin rendering. Unicode and offset contracts explain why a cleaned display string must still point back to the exact source bytes.

Generate several plausible forms

A candidate generator may combine an owner-approved alias table, a transliteration model and script-specific normalization. Record which process produced each form and its version. Candidate recall matters at this stage; precision comes later. A single top-ranked transliteration is a poor join key because vowels, aspirated consonants, abbreviations and regional conventions can create several spellings. Search by alias within an authorized entity catalog, then retain all plausible IDs until context is checked. Entity linking with a NIL decision supplies the broader candidate contract.

Control collision and access

Names can be homographs across scripts and companies can share a trading name. A normalized match can also cross a tenant boundary if the lookup table ignores scope. Restrict candidates to the caller’s permitted catalog before ranking. Require another independent clue such as a registered identifier, city or owner record for a consequential merge. Unknown or competing candidates should remain unresolved. The collision lesson turns this into a review decision.

Measure the downstream task

Evaluate candidate recall, unique-link precision, false merges and NIL decisions by script pair, entity type and spelling family. A single reference spelling can mark a valid alternate rendering wrong, so score the eventual entity lookup as well as the surface form. Split test data by entity family; otherwise a model may memorize the same organization under another spelling. Compare results against a source-script exact-match baseline. The project puts these checks into an intake workflow.

Implementation

python
def alias_candidates(source_name, source_script, permitted_ids, alias_rows):
    matches = []
    for alias in alias_rows:
        if alias["script"] != source_script:
            continue
        if alias["written_name"] != source_name:
            continue
        if alias["entity_id"] not in permitted_ids:
            continue
        matches.append({"entity_id": alias["entity_id"],
                        "alias_version": alias["version"]})
    return matches

aliases = [{"script": "Deva", "written_name": "सारिका",
            "entity_id": "vendor-47", "version": "alias-r4"},
           {"script": "Deva", "written_name": "सारिका",
            "entity_id": "vendor-82", "version": "alias-r4"}]
found = alias_candidates("सारिका", "Deva", {"vendor-47", "vendor-82"}, aliases)
assert {row["entity_id"] for row in found} == {"vendor-47", "vendor-82"}
assert alias_candidates("सारिका", "Deva", {"vendor-47"}, aliases) == [found[0]]

Performance and operating cost

The simple scan takes O(a) time and O(k) output space for a aliases and k permitted matches. An index on script and written form avoids scanning the whole catalog. This example only finds approved exact aliases; a transliteration model would propose additional forms, never authorize an entity merge.

Common Mistakes

  • Collapsing every spelling into a single irreversible canonical string.
  • Using the first Latin rendering as a primary key.
  • Ranking candidates before applying tenant scope.
  • Measuring only character similarity while ignoring false entity merges.

Read next

ai-data
natural-language-processing
Storage details