Skip to content
AITroveRead. Build. Understand.
Make this comfortable

PII detection and redaction on original text offsets

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A text pipeline should identify personal data before broad logging or model calls, but redaction must preserve auditability and avoid pretending detection is complete.

Define what requires protection

Names, emails, account numbers and addresses have different risks and formats. Start with a data inventory tied to the actual workflow, then specify which fields must be removed, tokenized or retained for an authorized purpose. A regex can find some structured values, while a trained span model may find contextual names and addresses. Neither can guarantee complete detection. Make a reviewed sample that includes punctuation, mixed scripts, OCR noise and copied signatures. Entity labels and exact-span metrics are useful foundations.

Detect before transforming

Run detectors against the captured original representation and produce typed half-open offsets. If normalization is needed for matching, maintain a verified mapping back to the original text. Merge overlapping detections under a written priority policy; otherwise replacing an inner email span first can invalidate an outer address span. Apply replacements from right to left so earlier offsets stay valid. Keep an irreversible public display separate from a restricted evidence store. Never write the raw value into ordinary metrics or error logs.

Choose a replacement contract

A fixed token such as [EMAIL] is suitable when exact recovery is unnecessary. Stable pseudonyms can preserve within-case joins, but they may reveal repeated identities across cases; scope them narrowly and protect the mapping. A replacement changes text length and all later offsets, so downstream annotations must refer either to the original or to a new mapped version. Annotate which representation every model consumes. Unicode boundaries matter when names contain combining marks.

Measure missed exposure

Report per-type recall, false-positive redactions, overlap conflicts and the volume sent for manual review. A single missed account number can matter more than several harmless over-redactions. Evaluate end-to-end exposure at every sink: analytics logs, search index, model input, generated summary and exported file. If the detector or mapping fails, fail closed for a sensitive sink. The audit path defines how to watch the result.

Implementation

python
def redact_original_text(message, spans):
    ordered = sorted(spans, key=lambda span: (span[0], span[1]))
    previous_end = 0
    for start, end, label in ordered:
        if not (0 <= start < end <= len(message)) or start < previous_end:
            raise ValueError("invalid or overlapping redaction spans")
        if not label.isidentifier():
            raise ValueError("invalid replacement label")
        previous_end = end
    redacted = message
    for start, end, label in reversed(ordered):
        redacted = redacted[:start] + f"[{label}]" + redacted[end:]
    return redacted

ticket = "Write to [email protected] about case 47."
assert redact_original_text(ticket, [(9, 24, "EMAIL")]) == "Write to [EMAIL] about case 47."

Performance and operating cost

Sorting s spans takes O(s log s). With immutable Python strings, repeated replacement can copy up to O(n·s) characters for an n-character message; a single streaming join reduces this to O(n + s) after sorting for high-volume use. Detection cost varies by regex and model. Report missed exposure on reviewed samples and sink-level leaks; latency alone cannot justify a weaker detector for sensitive text.

Common Mistakes

  • Assuming a regex or NER model finds every personal value.
  • Replacing left to right with offsets from the original text.
  • Logging raw detector inputs when an exception occurs.
  • Reusing original annotations against a redacted string without remapping.

Read next

Continue the workflow: Audit redacted text flows, retention and re-identification risk.

ai-data
natural-language-processing
Storage details