Text preprocessing is a versioned transform that must preserve characters relevant to labels, spans and serving behavior.
Unicode and tokenization: preserve meaning at the text boundary
Start with the task
Lowercasing, accent removal and punctuation deletion are choices, not automatic cleaning. An invoice code containing a hyphen may differ from one without it; negation can reverse a support request. Keep raw text and a versioned transformed representation. Test emoji, mixed scripts, composed accents, line breaks and ticket IDs before applying a global rule. The corpus contract] decides which changes preserve meaning.
Choose a token boundary
A word tokenizer, character n-grams and subword tokenizer create different vocabularies and unknown-token behavior. Record the tokenizer implementation, version, vocabulary and special-token map with the model. A model trained with one token map cannot safely serve another. For extraction tasks, keep a mapping from token positions back to original character offsets; normalized text may not have identical offsets.
Measure what is lost
Count empty outputs, unknown tokens and lengths before and after truncation by language or channel. A maximum length of 256 units can cut a long ticket before the actual request; inspect where the answer sits. When truncating, state whether you keep the beginning, end or selected segments. Span annotations] need the original text coordinate system.
Exercise edge cases
Use fixtures with a one-character ticket, a non-Latin message, combining marks and an unusually long message. Verify that each produces a deliberate output or a named rejection. Compare training and serving token IDs on the same raw bytes. A visual text match does not prove byte or token equality.
Implementation
import unicodedata
def normalize_ticket_text(raw_text):
if not isinstance(raw_text, str) or not raw_text.strip():
raise ValueError("empty ticket text")
normalized = unicodedata.normalize("NFC", raw_text)
normalized = " ".join(normalized.split())
if len(normalized) > 120_000:
raise ValueError("ticket exceeds processing limit")
return normalizedPerformance and operating cost
Normalization scans O(C) characters and creates O(C) output storage. Tokenization is usually linear in input length but vocabulary lookup and long-sequence model cost depend on the chosen architecture.
Common Mistakes
- Do not remove punctuation without checking task meaning.
- Do not apply a new tokenizer to an old fitted model.
- Do not interpret normalized offsets as raw-text offsets.
