A passage is a derived view of a particular source revision. A new boundary policy requires a new index and claim-link audit.
Passage revisions: chunk migrations, deletion and audit
Make passage identity reproducible
Derive a passage key from document ID, source revision, section path, original offsets and chunker version. Do not use passage text alone: repeated boilerplate may appear in several locations. Store a digest of the source region for validation. A citation to a passage must resolve to the same text and access scope seen during answer generation. Section structure defines the source coordinates.
Stage a new chunk policy
A new tokenizer, overlap or heading parser can reassign passage boundaries across the corpus. Build a shadow index, compare document and passage counts, check supporting-passage recall and inspect citations from cached answers. Keep old and new indices separate until validated. Do not let an answer cite an old passage ID against a new source revision; the linked claim may no longer be supported.
Propagate deletion and access
Removing a document must remove every passage, vector, keyword posting and cached answer tied to it. Scope checks should use the current source policy, not a stale access flag copied onto a passage. An index that serves a deleted document during rollback is not a safe rollback. The privacy pipeline handles the broader derivative inventory.
Measure migration risk
Report stale passage references, orphaned citations, wrong-section hits, missing deletions and index parity on exact identifiers. Test a document whose heading moved while body text stayed the same, as well as one whose warning changed. The runbook project makes passage migration a release gate before moving the search alias.
Implementation
import hashlib
def passage_identity(document_id, revision, chunker_version,
section_path, start, end):
if not document_id or not revision or not 0 <= start < end:
raise ValueError("invalid passage identity")
parts = (document_id, revision, chunker_version,
"/".join(section_path), str(start), str(end))
return hashlib.sha256("|".join(parts).encode("utf-8")).hexdigest()
first = passage_identity("rb-47", "r8", "chunk-v3", ["Recovery"], 0, 47)
second = passage_identity("rb-47", "r9", "chunk-v3", ["Recovery"], 0, 47)
assert first != second
Performance and operating cost
Hashing the identity fields costs O(k) for their combined byte length and produces fixed-size keys. A shadow reindex costs O(total document text) plus storage writes; citation replay adds query work. Pay that temporary cost before promotion, because stale citations can make a fluent answer appear grounded in text that is no longer current.
Common Mistakes
- Using passage text alone as a stable ID.
- Serving cached answers after their supporting revision changed.
- Leaving deleted passages in the rollback index.
- Comparing only document counts while section boundaries shifted.
Read next
- Long documents: section boundaries and passage provenance
- Project: index runbook sections with revision-safe passages
- Text embeddings: pair labels, hard negatives and versioned vectors
- Audit redacted text flows, retention and re-identification risk
- Generated answers: claim evidence and unanswered state
Continue the workflow: Abbreviation registry: ownership, versions and index migration.
Continue the workflow: Semantic document diffs: changed facts versus wording edits.
Continue the workflow: Retrieval positives: query intent, passage identity and revision.
