Accept multilingual support text, keep source spans stable and review lookalike service identifiers before they enter routing or search.
Project: preserve text identity in support intake
Set distinct field contracts
A support message is free-form multilingual text; a service slug is a controlled identifier. Store the original message and identifier separately. Build an NFC comparison key for search, but retain source offsets for highlights and audits. Apply a field-specific identifier policy only to the slug. The service should never rewrite a customer quotation merely to make search easier. Normalization and offsets define the source-to-key relationship.
Build an adversarial review set
Include canonically equivalent names, combining marks, full-width text, mixed-script lookalikes, ordinary code switching and copied screenshots that passed through OCR. Annotate which field is an identifier, which is prose and what the user should see. Group near-identical tickets and document revisions in the same evaluation split. Include a suspicious service slug whose visual form resembles an approved one, plus legitimate multilingual prose that must pass without identifier alarms.
Stage routing and review
Normalize comparison keys without discarding originals, then run field-specific identifier triage. Hold suspicious new slugs for owner review; do not auto-merge them with the nearest existing slug. Continue processing the free-form message under ordinary privacy and language rules. Any extracted span shown to a reviewer must map back to the original source revision. The identifier policy is a review gate, not a general language filter.
Release against user-visible errors
Measure wrong routing, missed suspicious slugs, false alarms by script, broken highlights and search-result changes. Inspect whether a customer can still find a document using an alternate Unicode form while the displayed title retains its source spelling. Shadow the policy before release and keep old index keys until migration passes. Roll back if legitimate messages are blocked or if identifiers collapse into the same authority key without review.
Implementation
import unicodedata
def stage_intake(message, service_slug, allowed_slugs):
search_key = unicodedata.normalize("NFC", message)
if service_slug not in allowed_slugs:
return {"state": "review-identifier", "original_message": message,
"search_key": search_key, "service_slug": service_slug}
return {"state": "accepted", "original_message": message,
"search_key": search_key, "service_slug": service_slug}
record = stage_intake("café gateway alert", "gateway-west",
{"gateway-west", "worker-east"})
assert record["state"] == "accepted"
assert record["original_message"] != record["search_key"]
assert stage_intake("gateway alert", "gаteway-west", {"gateway-west"})["state"] == "review-identifier"
Performance and operating cost
Creating a comparison key for n characters is O(n) time and space; membership in an in-memory allowed-slug set is expected O(1). The code uses exact allowlisting for illustration, not a full confusable algorithm. A production registration path also needs a versioned security mapping and owner review. Keep search recall and identifier authority as separate measured outcomes.
Common Mistakes
- Normalizing away the only stored copy of the message.
- Applying identifier rejection rules to all customer prose.
- Auto-merging a suspicious slug with a visually similar approved one.
- Highlighting the normalized key with offsets from the original message.
