A chunk is a retrievable unit with text, document revision and structural context; its boundaries affect both recall and evidence quality.
Chunking documents: preserve section identity and answer boundaries
Split on meaning first
A policy heading and its exception should remain connected. Blind fixed-character windows may separate the rule from a qualifying sentence and cause a plausible but wrong answer. Prefer section or paragraph boundaries, then enforce a size limit. Carry heading path, document ID, revision and exact offsets. Document identity] remains attached to every chunk.
Budget overlap deliberately
Small chunks improve pinpoint retrieval but can lose context. Large chunks carry context but consume more prompt budget and may rank by unrelated words. Overlap helps at boundaries but creates near duplicates. Count distinct parent documents in the final context so four overlapping chunks do not crowd out another necessary policy.
Preserve reproducibility
Version parser, normalization, chunking parameters and embedding model together. If extraction from a PDF changes line order, a stable document revision can still produce different chunks. Store a content hash and offset or section identifier for each passage. An answer audit should resolve exactly the passage the model saw, not today’s edited page.
Test difficult structure
Create a table-like policy with a footnote, a long section and two near-identical headings. Check that the exception stays retrievable with the rule. Verify each chunk respects its input limit and can be mapped back to an exact document revision. Retrieval evaluation] will show whether the chosen boundaries help real questions.
Implementation
import hashlib
import json
def make_chunk_id(document_id, revision, section_path, text):
identity = json.dumps([document_id, revision, section_path, text], ensure_ascii=False)
return hashlib.sha256(identity.encode("utf-8")).hexdigest()[:24]
chunk_record = {"chunk_id": make_chunk_id(document_id, revision, section_path, passage_text),
"document_id": document_id, "revision": revision,
"section": section_path, "text": passage_text}Performance and operating cost
Parsing and splitting C characters is at least O(C); embedding every chunk adds model cost proportional to chunk count and length. Overlap raises both index size and query-context duplication.
Common Mistakes
- Do not separate an exception from the rule it qualifies.
- Do not lose document revision and offsets when storing chunks.
- Do not let overlapping chunks monopolize the context.
