A visually familiar service name can contain different characters. Triage suspicious identifiers without treating every multilingual phrase as an attack.
Confusable identifiers: script policy and review queues
Scope the policy to identifiers
An incident title may mix scripts legitimately. A machine-readable service slug, however, usually has a narrower alphabet contract. Define allowed characters and case policy for each identifier field, rather than applying one mixed-script rejection rule to all natural language. Preserve the original bytes and code points for audit. Script profiling helps describe a message; a security decision needs field-specific rules and a reviewed allowlist.
Separate normalization from lookalikes
Canonical normalization does not make every visually similar character equal. A Latin letter and a similar-looking letter from another script may remain different keys. A full confusable analysis needs a versioned character mapping and a declared supported repertoire; a small script check is only a triage signal. Do not use a generated visual skeleton as the displayed identifier or as unquestioned proof of identity. Normalization policy provides a separate comparison step.
Make collisions reviewable
When a new identifier resembles an existing service slug, stage both originals, code-point names, owner and registry IDs for review. Restrict creation or require approval according to the field policy. Do not silently merge two accounts or services because a visual mapping matches. Conversely, avoid flagging every multilingual customer message; the policy belongs to identifier creation and high-risk routing, not ordinary prose. Keep a decision record so a later policy revision can revisit accepted exceptions.
Measure both safety and friction
Audit missed lookalikes, false alarms by language, reviewer time and accidental collisions after registry changes. Use realistic font renderings in manual review because visual similarity depends on display context. A triage rule that blocks legitimate names can harm users even while catching a few suspicious cases. The intake project tests how the identifier rule interacts with search and customer text.
Implementation
import unicodedata
def identifier_script_hint(identifier):
scripts = set()
for character in identifier:
if character.isalpha():
name = unicodedata.name(character, "")
scripts.add(name.split(" ", 1)[0] if name else "UNKNOWN")
return {"scripts": sorted(scripts), "review": len(scripts) > 1}
latin = identifier_script_hint("gateway47")
mixed = identifier_script_hint("gаteway47") # second character is Cyrillic
assert latin["scripts"] == ["LATIN"]
assert mixed["review"]
Performance and operating cost
This illustrative name-based scan is O(n) time and O(s) space for n characters and s observed name prefixes. It is a triage hint, not complete script detection or confusable comparison: common, inherited and complex script properties need a dedicated, versioned standard-data implementation. Apply the stronger check at identifier boundaries, where review volume is manageable.
Common Mistakes
- Treating a simple mixed-script hint as a complete confusable detector.
- Rejecting multilingual prose under an identifier-only policy.
- Displaying a transformed skeleton instead of the submitted spelling.
- Automatically merging records because two identifiers look alike.
