Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Long documents: evidence indexing and section coverage

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A long runbook can hold the answer across distant sections. Retrieval needs stable passage identity and a way to notice missing evidence.

Index passages with their place in the document

Split a runbook by meaningful headings, steps and revision boundaries rather than a blind character count alone. Store document ID, section path, passage index, source offsets, revision and access scope. A passage about a rollback may depend on a precondition several sections earlier; the passage index must retain that parent-child relation. Section boundaries define the initial cut, while revision auditing detects stale indexed copies.

Retrieve for each required subquestion

“How do we roll back and verify recovery?” asks for at least two kinds of evidence. A single high-scoring passage may explain rollback but omit verification. Decompose the information need into required facets, retrieve candidates for each, and record which facets remain uncovered. Do not hide uncovered requirements by letting one verbose passage dominate the score. Query provenance keeps each reformulated subquestion tied to the original operator request.

Preserve evidence order and access

Procedures often have prerequisites, actions and checks. Reordering retrieved passages can turn a safe procedure into an unsafe one. Keep section order and explicit references when assembling evidence. Apply access controls to every passage before use; a public answer cannot quote a restricted appendix merely because it completes a missing facet. If an essential facet is restricted or absent, return a partial result with that limit stated, rather than filling the gap from model memory.

Audit coverage separately from relevance

Measure retrieval recall for each required facet, section-order mistakes, stale passages and unsupported final claims. A top-k score may look good while the only verification step sits below the cutoff. Include long runbooks with repeated headings, appendix corrections and multiple revisions. Reviewers need the exact passage path and offset for each answer sentence. The long-runbook project evaluates this complete workflow.

Implementation

python
def uncovered_facets(required_facets, passages):
    covered = set()
    for passage in passages:
        covered.update(passage["facets"])
    return sorted(set(required_facets) - covered)

selected = [{"passage_id": "rollback-47", "facets": {"rollback"}},
            {"passage_id": "checks-82", "facets": {"verification"}}]
assert uncovered_facets({"rollback", "verification"}, selected) == []
assert uncovered_facets({"rollback", "verification", "prerequisite"}, selected) == ["prerequisite"]

Performance and operating cost

Coverage accounting takes O(p + f) expected time for p passage-facet memberships and f required facets, plus O(f) output space. Retrieval over a long corpus adds index and ranking costs. More passages do not guarantee complete evidence; measure facet recall and access-filtered availability rather than only retrieval latency.

Common Mistakes

  • Using one relevant paragraph as proof that every subquestion is answered.
  • Losing prerequisite order while concatenating passages.
  • Dropping source revisions from indexed passage IDs.
  • Treating a restricted passage as usable evidence for every reader.

Read next

Continue the workflow: Multi-step questions: decompose answer slots and dependencies.

ai-data
natural-language-processing
Storage details