Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: preserve text identity in support intake

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Accept multilingual support text, keep source spans stable and review lookalike service identifiers before they enter routing or search.

Set distinct field contracts

A support message is free-form multilingual text; a service slug is a controlled identifier. Store the original message and identifier separately. Build an NFC comparison key for search, but retain source offsets for highlights and audits. Apply a field-specific identifier policy only to the slug. The service should never rewrite a customer quotation merely to make search easier. Normalization and offsets define the source-to-key relationship.

Build an adversarial review set

Include canonically equivalent names, combining marks, full-width text, mixed-script lookalikes, ordinary code switching and copied screenshots that passed through OCR. Annotate which field is an identifier, which is prose and what the user should see. Group near-identical tickets and document revisions in the same evaluation split. Include a suspicious service slug whose visual form resembles an approved one, plus legitimate multilingual prose that must pass without identifier alarms.

Stage routing and review

Normalize comparison keys without discarding originals, then run field-specific identifier triage. Hold suspicious new slugs for owner review; do not auto-merge them with the nearest existing slug. Continue processing the free-form message under ordinary privacy and language rules. Any extracted span shown to a reviewer must map back to the original source revision. The identifier policy is a review gate, not a general language filter.

Release against user-visible errors

Measure wrong routing, missed suspicious slugs, false alarms by script, broken highlights and search-result changes. Inspect whether a customer can still find a document using an alternate Unicode form while the displayed title retains its source spelling. Shadow the policy before release and keep old index keys until migration passes. Roll back if legitimate messages are blocked or if identifiers collapse into the same authority key without review.

Implementation

python
import unicodedata

def stage_intake(message, service_slug, allowed_slugs):
    search_key = unicodedata.normalize("NFC", message)
    if service_slug not in allowed_slugs:
        return {"state": "review-identifier", "original_message": message,
                "search_key": search_key, "service_slug": service_slug}
    return {"state": "accepted", "original_message": message,
            "search_key": search_key, "service_slug": service_slug}

record = stage_intake("café gateway alert", "gateway-west",
                      {"gateway-west", "worker-east"})
assert record["state"] == "accepted"
assert record["original_message"] != record["search_key"]
assert stage_intake("gateway alert", "gаteway-west", {"gateway-west"})["state"] == "review-identifier"

Performance and operating cost

Creating a comparison key for n characters is O(n) time and space; membership in an in-memory allowed-slug set is expected O(1). The code uses exact allowlisting for illustration, not a full confusable algorithm. A production registration path also needs a versioned security mapping and owner review. Keep search recall and identifier authority as separate measured outcomes.

Common Mistakes

  • Normalizing away the only stored copy of the message.
  • Applying identifier rejection rules to all customer prose.
  • Auto-merging a suspicious slug with a visually similar approved one.
  • Highlighting the normalized key with offsets from the original message.

Read next

ai-data
natural-language-processing
Storage details