Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Near-duplicate text families: detect overlap without erasing meaning

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A corpus can contain revised, quoted and templated copies of one document. Group likely relatives before measuring generalization, while keeping distinct claims separate.

Define the family, not just the string

An incident ticket may be copied into a handoff, amended after a deployment and quoted in a later support reply. Exact hashes find identical bytes but miss these relatives. Store document ID, origin ID, revision ID, quoted source spans, capture time and access policy. A text-family ID means that examples share enough origin or content that a model could memorize one while being tested on another. It does not mean their labels or factual claims are equal. Corpus identity supplies the provenance that a similarity score cannot recover.

Generate candidates conservatively

Normalize only for candidate discovery. Token shingles and Jaccard overlap can flag edited copies, while exact hashes cheaply identify identical text. Keep the original bytes and offsets for review. A shared footer can make unrelated tickets look alike; a changed negation can make nearly identical tickets mean opposite things. Compare substantive spans, templates and source lineage before merging families. Paraphrase review asks whether meaning is preserved, a different question from whether two records should be separated during evaluation.

Review connected components

Pairwise similarity is not transitive. If ticket A resembles B and B resembles C, a connected-component rule may group A with C even when their claims conflict. Use components as split-safety candidates, record the edges that joined them and review unusually large or mixed-label groups. When lineage is known, prefer it to a text threshold. A family can include conflicting versions because the split rule prevents memorization; the label policy must still evaluate each revision on its own terms.

Measure the audit itself

Report exact duplicates, candidate near-duplicates, reviewer-confirmed families and cross-split family overlap. Slice by source, language, template and date. Inspect false joins, especially boilerplate-heavy documents, and missed joins after OCR or translation. Family-aware splitting uses the audited IDs; the project turns them into a release gate.

Implementation

python
import re

def word_shingles(document_text, width=3):
    if width < 1:
        raise ValueError("width must be positive")
    tokens = re.findall(r"\w+", document_text.casefold())
    if not tokens:
        return set()
    if len(tokens) < width:
        return {tuple(tokens)}
    return {tuple(tokens[index:index + width])
            for index in range(len(tokens) - width + 1)}

def shingle_overlap(left_text, right_text):
    left = word_shingles(left_text)
    right = word_shingles(right_text)
    if not left or not right:
        return 0.0
    return len(left & right) / len(left | right)

original = "Gateway west rejected retry for request 47 after timeout"
revision = "Gateway west rejected retry for request 47 after timeout again"
unrelated = "Billing north confirmed refund for invoice 82"
assert shingle_overlap(original, revision) > shingle_overlap(original, unrelated)
assert shingle_overlap("", revision) == 0.0

Performance and operating cost

Building shingles takes O(n) time and space for n tokens. Comparing two shingle sets takes expected O(n + m) time and space, but comparing every pair in a corpus is O(d²) for d documents before token cost. Block by lineage, service or a candidate index at scale. This lexical signal neither proves shared origin nor licenses deletion of a contradictory revision.

Common Mistakes

  • Treating an exact hash as a complete near-duplicate audit.
  • Collapsing two revisions into one label when their instructions conflict.
  • Joining records only because they share a long template footer.
  • Discarding original text and offsets after normalization.

Read next

ai-data
natural-language-processing
Storage details