Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Passage revisions: chunk migrations, deletion and audit

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A passage is a derived view of a particular source revision. A new boundary policy requires a new index and claim-link audit.

Make passage identity reproducible

Derive a passage key from document ID, source revision, section path, original offsets and chunker version. Do not use passage text alone: repeated boilerplate may appear in several locations. Store a digest of the source region for validation. A citation to a passage must resolve to the same text and access scope seen during answer generation. Section structure defines the source coordinates.

Stage a new chunk policy

A new tokenizer, overlap or heading parser can reassign passage boundaries across the corpus. Build a shadow index, compare document and passage counts, check supporting-passage recall and inspect citations from cached answers. Keep old and new indices separate until validated. Do not let an answer cite an old passage ID against a new source revision; the linked claim may no longer be supported.

Propagate deletion and access

Removing a document must remove every passage, vector, keyword posting and cached answer tied to it. Scope checks should use the current source policy, not a stale access flag copied onto a passage. An index that serves a deleted document during rollback is not a safe rollback. The privacy pipeline handles the broader derivative inventory.

Measure migration risk

Report stale passage references, orphaned citations, wrong-section hits, missing deletions and index parity on exact identifiers. Test a document whose heading moved while body text stayed the same, as well as one whose warning changed. The runbook project makes passage migration a release gate before moving the search alias.

Implementation

python
import hashlib

def passage_identity(document_id, revision, chunker_version,
                     section_path, start, end):
    if not document_id or not revision or not 0 <= start < end:
        raise ValueError("invalid passage identity")
    parts = (document_id, revision, chunker_version,
             "/".join(section_path), str(start), str(end))
    return hashlib.sha256("|".join(parts).encode("utf-8")).hexdigest()

first = passage_identity("rb-47", "r8", "chunk-v3", ["Recovery"], 0, 47)
second = passage_identity("rb-47", "r9", "chunk-v3", ["Recovery"], 0, 47)
assert first != second

Performance and operating cost

Hashing the identity fields costs O(k) for their combined byte length and produces fixed-size keys. A shadow reindex costs O(total document text) plus storage writes; citation replay adds query work. Pay that temporary cost before promotion, because stale citations can make a fluent answer appear grounded in text that is no longer current.

Common Mistakes

  • Using passage text alone as a stable ID.
  • Serving cached answers after their supporting revision changed.
  • Leaving deleted passages in the rollback index.
  • Comparing only document counts while section boundaries shifted.

Read next

Continue the workflow: Abbreviation registry: ownership, versions and index migration.

Continue the workflow: Semantic document diffs: changed facts versus wording edits.

Continue the workflow: Retrieval positives: query intent, passage identity and revision.

ai-data
natural-language-processing
Storage details