Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Chunking documents: preserve section identity and answer boundaries

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A chunk is a retrievable unit with text, document revision and structural context; its boundaries affect both recall and evidence quality.

Split on meaning first

A policy heading and its exception should remain connected. Blind fixed-character windows may separate the rule from a qualifying sentence and cause a plausible but wrong answer. Prefer section or paragraph boundaries, then enforce a size limit. Carry heading path, document ID, revision and exact offsets. Document identity] remains attached to every chunk.

Budget overlap deliberately

Small chunks improve pinpoint retrieval but can lose context. Large chunks carry context but consume more prompt budget and may rank by unrelated words. Overlap helps at boundaries but creates near duplicates. Count distinct parent documents in the final context so four overlapping chunks do not crowd out another necessary policy.

Preserve reproducibility

Version parser, normalization, chunking parameters and embedding model together. If extraction from a PDF changes line order, a stable document revision can still produce different chunks. Store a content hash and offset or section identifier for each passage. An answer audit should resolve exactly the passage the model saw, not today’s edited page.

Test difficult structure

Create a table-like policy with a footnote, a long section and two near-identical headings. Check that the exception stays retrievable with the rule. Verify each chunk respects its input limit and can be mapped back to an exact document revision. Retrieval evaluation] will show whether the chosen boundaries help real questions.

Implementation

python
import hashlib
import json

def make_chunk_id(document_id, revision, section_path, text):
    identity = json.dumps([document_id, revision, section_path, text], ensure_ascii=False)
    return hashlib.sha256(identity.encode("utf-8")).hexdigest()[:24]

chunk_record = {"chunk_id": make_chunk_id(document_id, revision, section_path, passage_text),
                "document_id": document_id, "revision": revision,
                "section": section_path, "text": passage_text}

Performance and operating cost

Parsing and splitting C characters is at least O(C); embedding every chunk adds model cost proportional to chunk count and length. Overlap raises both index size and query-context duplication.

Common Mistakes

  • Do not separate an exception from the rule it qualifies.
  • Do not lose document revision and offsets when storing chunks.
  • Do not let overlapping chunks monopolize the context.

Read next

ai-data
retrieval-ai
Storage details