A corpus can contain revised, quoted and templated copies of one document. Group likely relatives before measuring generalization, while keeping distinct claims separate.
Near-duplicate text families: detect overlap without erasing meaning
Define the family, not just the string
An incident ticket may be copied into a handoff, amended after a deployment and quoted in a later support reply. Exact hashes find identical bytes but miss these relatives. Store document ID, origin ID, revision ID, quoted source spans, capture time and access policy. A text-family ID means that examples share enough origin or content that a model could memorize one while being tested on another. It does not mean their labels or factual claims are equal. Corpus identity supplies the provenance that a similarity score cannot recover.
Generate candidates conservatively
Normalize only for candidate discovery. Token shingles and Jaccard overlap can flag edited copies, while exact hashes cheaply identify identical text. Keep the original bytes and offsets for review. A shared footer can make unrelated tickets look alike; a changed negation can make nearly identical tickets mean opposite things. Compare substantive spans, templates and source lineage before merging families. Paraphrase review asks whether meaning is preserved, a different question from whether two records should be separated during evaluation.
Review connected components
Pairwise similarity is not transitive. If ticket A resembles B and B resembles C, a connected-component rule may group A with C even when their claims conflict. Use components as split-safety candidates, record the edges that joined them and review unusually large or mixed-label groups. When lineage is known, prefer it to a text threshold. A family can include conflicting versions because the split rule prevents memorization; the label policy must still evaluate each revision on its own terms.
Measure the audit itself
Report exact duplicates, candidate near-duplicates, reviewer-confirmed families and cross-split family overlap. Slice by source, language, template and date. Inspect false joins, especially boilerplate-heavy documents, and missed joins after OCR or translation. Family-aware splitting uses the audited IDs; the project turns them into a release gate.
Implementation
import re
def word_shingles(document_text, width=3):
if width < 1:
raise ValueError("width must be positive")
tokens = re.findall(r"\w+", document_text.casefold())
if not tokens:
return set()
if len(tokens) < width:
return {tuple(tokens)}
return {tuple(tokens[index:index + width])
for index in range(len(tokens) - width + 1)}
def shingle_overlap(left_text, right_text):
left = word_shingles(left_text)
right = word_shingles(right_text)
if not left or not right:
return 0.0
return len(left & right) / len(left | right)
original = "Gateway west rejected retry for request 47 after timeout"
revision = "Gateway west rejected retry for request 47 after timeout again"
unrelated = "Billing north confirmed refund for invoice 82"
assert shingle_overlap(original, revision) > shingle_overlap(original, unrelated)
assert shingle_overlap("", revision) == 0.0
Performance and operating cost
Building shingles takes O(n) time and space for n tokens. Comparing two shingle sets takes expected O(n + m) time and space, but comparing every pair in a corpus is O(d²) for d documents before token cost. Block by lineage, service or a candidate index at scale. This lexical signal neither proves shared origin nor licenses deletion of a contradictory revision.
Common Mistakes
- Treating an exact hash as a complete near-duplicate audit.
- Collapsing two revisions into one label when their instructions conflict.
- Joining records only because they share a long template footer.
- Discarding original text and offsets after normalization.
Read next
- Family-aware text splits and evaluation contamination
- Project: audit ticket families before publishing a text classifier score
- Text corpus contracts: identity, label timing and annotation rules
- Duplicate clusters: review transitivity, bridges and source identity
- Semantic document diffs: changed facts versus wording edits
