Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Unicode normalization: search keys without losing source offsets

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Canonical and compatibility normalization solve different comparison problems. Keep original text and map transformed spans back to it.

Choose normalization by field

Two visually similar strings may use different code-point sequences. NFC can make canonically equivalent forms comparable; compatibility normalization may also fold distinctions that a document preserves intentionally. Use a documented comparison form for each field: a search key, identifier and quoted source passage may need different treatment. Do not turn the normalized value into the only stored copy. Tokenization boundaries depend on the chosen text form and must be versioned with the index.

Retain an offset map

A detected entity span on transformed text is not automatically an offset into the original text. Normalization can change code-point count, and case folding can do so too. Record source revision, normalization policy and a mapping from transformed ranges to original ranges, or run span extraction on the original and normalize only comparison keys. A highlight that points into the wrong customer name is a real product error, not a cosmetic discrepancy. Span annotations must stay anchored to the original.

Keep search and identity separate

Search may index alternate forms to improve recall, while an account identifier or document ID needs its own exact equality and security policy. Do not use a broad compatibility fold as an implicit authority to merge accounts, incidents or entities. Show the user the original spelling, and make the matching policy inspectable. Confusable triage addresses lookalike characters that ordinary normalization does not necessarily collapse.

Audit language and revision effects

Test combining marks, full-width forms, ligatures, emoji sequences and text from each supported script. Measure retrieved-result changes, broken highlights and accidental key collisions. Reindex when a normalization policy changes, and retain the old policy version for comparison. A transformation that improves one language may damage another or convert two distinct operational IDs into the same key. The intake project keeps those cases visible during release review.

Implementation

python
import unicodedata

def comparison_record(source_text, policy_version):
    return {"original": source_text,
            "nfc_key": unicodedata.normalize("NFC", source_text),
            "policy_version": policy_version}

composed = comparison_record("café", "norm-r4")
decomposed = comparison_record("café", "norm-r4")
assert composed["nfc_key"] == decomposed["nfc_key"]
assert composed["original"] != decomposed["original"]

Performance and operating cost

Normalizing n code points is O(n) time and O(n) output space. Maintaining a reliable transformed-to-source offset map also costs O(n) and needs explicit tests for expansions and contractions. The short code only builds a comparison record; it deliberately does not claim that offsets in the key can be reused on the original.

Common Mistakes

  • Overwriting source text with a normalized key.
  • Using transformed offsets to highlight the original passage.
  • Applying compatibility folding to authoritative IDs without a collision policy.
  • Changing normalization without rebuilding affected search indexes.

Read next

ai-data
natural-language-processing
Storage details