Canonical and compatibility normalization solve different comparison problems. Keep original text and map transformed spans back to it.
Unicode normalization: search keys without losing source offsets
Choose normalization by field
Two visually similar strings may use different code-point sequences. NFC can make canonically equivalent forms comparable; compatibility normalization may also fold distinctions that a document preserves intentionally. Use a documented comparison form for each field: a search key, identifier and quoted source passage may need different treatment. Do not turn the normalized value into the only stored copy. Tokenization boundaries depend on the chosen text form and must be versioned with the index.
Retain an offset map
A detected entity span on transformed text is not automatically an offset into the original text. Normalization can change code-point count, and case folding can do so too. Record source revision, normalization policy and a mapping from transformed ranges to original ranges, or run span extraction on the original and normalize only comparison keys. A highlight that points into the wrong customer name is a real product error, not a cosmetic discrepancy. Span annotations must stay anchored to the original.
Keep search and identity separate
Search may index alternate forms to improve recall, while an account identifier or document ID needs its own exact equality and security policy. Do not use a broad compatibility fold as an implicit authority to merge accounts, incidents or entities. Show the user the original spelling, and make the matching policy inspectable. Confusable triage addresses lookalike characters that ordinary normalization does not necessarily collapse.
Audit language and revision effects
Test combining marks, full-width forms, ligatures, emoji sequences and text from each supported script. Measure retrieved-result changes, broken highlights and accidental key collisions. Reindex when a normalization policy changes, and retain the old policy version for comparison. A transformation that improves one language may damage another or convert two distinct operational IDs into the same key. The intake project keeps those cases visible during release review.
Implementation
import unicodedata
def comparison_record(source_text, policy_version):
return {"original": source_text,
"nfc_key": unicodedata.normalize("NFC", source_text),
"policy_version": policy_version}
composed = comparison_record("café", "norm-r4")
decomposed = comparison_record("café", "norm-r4")
assert composed["nfc_key"] == decomposed["nfc_key"]
assert composed["original"] != decomposed["original"]
Performance and operating cost
Normalizing n code points is O(n) time and O(n) output space. Maintaining a reliable transformed-to-source offset map also costs O(n) and needs explicit tests for expansions and contractions. The short code only builds a comparison record; it deliberately does not claim that offsets in the key can be reused on the original.
Common Mistakes
- Overwriting source text with a normalized key.
- Using transformed offsets to highlight the original passage.
- Applying compatibility folding to authoritative IDs without a collision policy.
- Changing normalization without rebuilding affected search indexes.
