A fixed-length chunk can split a warning from its condition. Preserve headings, offsets and parent revisions while selecting passages.
Long documents: section boundaries and passage provenance
Start with document structure
Represent title, heading hierarchy, paragraph, list and table separately before creating model-sized passages. A runbook’s “Do not retry” warning may govern the next three steps; splitting at an arbitrary character count can strand the warning outside the retrieved chunk. Keep the parent document ID, section path and original offsets for every passage. OCR geometry adds page coordinates when the source is scanned.
Choose a boundary policy
Prefer heading and paragraph boundaries, then split oversized sections by sentences with controlled overlap. The overlap helps retrieval but creates duplicates in search results and summaries. Do not move a table header away from its rows or a code block away from its explanation. Record the chunker version and exact source revision; a changed heading can alter many passage IDs even when most sentences remain.
Preserve meaning in retrieval
Index a passage with enough parent context to interpret its conditions, but do not add unrelated neighboring text just to fill a token budget. Reconstruct a result with its heading and source link. Compare retrieval at passage and document level: a correct document with the wrong section is still a bad answer. Ranking and access filters must apply to the source before a passage is returned.
Evaluate boundary errors
Create questions whose answer crosses a paragraph break, lives in a table or depends on a nearby warning. Report exact supporting-passage recall, wrong-section rate, duplicate result rate and answer claim errors. Test deleted and revised documents. Revision audits manage derived IDs; the project tests the full search path.
Implementation
def passage_from_section(document_id, revision, section, start, end):
content = section["text"]
if not 0 <= start < end <= len(content):
raise ValueError("passage outside section")
return {"document_id": document_id, "revision": revision,
"section_path": tuple(section["path"]),
"start": start, "end": end, "text": content[start:end]}
section = {"path": ["Recovery", "Retry policy"],
"text": "Do not retry until the queue drains."}
passage = passage_from_section("rb-47", "r8", section, 0, len(section["text"]))
assert passage["section_path"] == ("Recovery", "Retry policy")
Performance and operating cost
Slicing a passage copies O(length of passage) characters, while segmenting the full document is O(n) for n characters under a linear boundary scan. Overlap increases index bytes and candidate counts. Measure answer quality against those costs; the shortest chunks are not always best if they lose the condition that makes an instruction safe.
Common Mistakes
- Splitting a warning from the procedure it qualifies.
- Indexing table rows without their headers.
- Treating correct-document retrieval as correct-passage retrieval.
- Reusing passage IDs after the parent document changes.
Read next
- Passage revisions: chunk migrations, deletion and audit
- Project: index runbook sections with revision-safe passages
- Hybrid text ranking with access filters and reranking
- OCR text identity: geometry, reading order and offsets
- Generated answers: claim evidence and unanswered state
Continue the workflow: Sense decisions: context windows, policy versions and review.
Continue the workflow: Abbreviations: definition scope, collisions and unknown forms.
Continue the workflow: Long documents: evidence indexing and section coverage.
Continue the workflow: Email threads: quote boundaries and authored-text attribution.
