Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Unicode and tokenization: preserve meaning at the text boundary

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Text preprocessing is a versioned transform that must preserve characters relevant to labels, spans and serving behavior.

Start with the task

Lowercasing, accent removal and punctuation deletion are choices, not automatic cleaning. An invoice code containing a hyphen may differ from one without it; negation can reverse a support request. Keep raw text and a versioned transformed representation. Test emoji, mixed scripts, composed accents, line breaks and ticket IDs before applying a global rule. The corpus contract] decides which changes preserve meaning.

Choose a token boundary

A word tokenizer, character n-grams and subword tokenizer create different vocabularies and unknown-token behavior. Record the tokenizer implementation, version, vocabulary and special-token map with the model. A model trained with one token map cannot safely serve another. For extraction tasks, keep a mapping from token positions back to original character offsets; normalized text may not have identical offsets.

Measure what is lost

Count empty outputs, unknown tokens and lengths before and after truncation by language or channel. A maximum length of 256 units can cut a long ticket before the actual request; inspect where the answer sits. When truncating, state whether you keep the beginning, end or selected segments. Span annotations] need the original text coordinate system.

Exercise edge cases

Use fixtures with a one-character ticket, a non-Latin message, combining marks and an unusually long message. Verify that each produces a deliberate output or a named rejection. Compare training and serving token IDs on the same raw bytes. A visual text match does not prove byte or token equality.

Implementation

python
import unicodedata

def normalize_ticket_text(raw_text):
    if not isinstance(raw_text, str) or not raw_text.strip():
        raise ValueError("empty ticket text")
    normalized = unicodedata.normalize("NFC", raw_text)
    normalized = " ".join(normalized.split())
    if len(normalized) > 120_000:
        raise ValueError("ticket exceeds processing limit")
    return normalized

Performance and operating cost

Normalization scans O(C) characters and creates O(C) output storage. Tokenization is usually linear in input length but vocabulary lookup and long-sequence model cost depend on the chosen architecture.

Common Mistakes

  • Do not remove punctuation without checking task meaning.
  • Do not apply a new tokenizer to an old fitted model.
  • Do not interpret normalized offsets as raw-text offsets.

Read next

ai-data
natural-language-processing
Storage details