Retrieve prerequisites, actions and verification steps from a long runbook, then release only answer sentences with active evidence.
Project: answer long-runbook questions with complete evidence
Set the operational question
An engineer asks how to roll back gateway-west in production and verify recovery. The answer needs a prerequisite, rollback action and verification step. Treat each as a required facet, not a reason to generate a long generic response. Every sentence points to a current, access-permitted passage. The system must say which facet is missing if the runbook does not contain it. Coverage accounting exposes those gaps before generation.
Create a difficult holdout
Use long runbooks with repeated headings, staging and production variants, superseded appendices, restricted sections and steps separated by many pages. Group all revisions of the same runbook in one split. Reviewers label exact supporting spans and whether a passage is current for the requested environment. Include questions that the runbook cannot answer, plus cases where two active sections conflict. Claim alignment handles those disputes.
Stage retrieval and synthesis
Retrieve candidates by facet under the reader’s access scope. Validate document revision, environment and section order. Build an evidence ledger listing the proposed answer sentence, supporting passage IDs and condition for each claim. If a facet is uncovered or two instructions conflict, route the draft to review instead of manufacturing a confident procedure. Do not add a step from a nearby but different service merely to make the response complete.
Gate the answer people see
Measure facet recall, wrong-environment steps, unsupported sentences, missed conflicts, stale evidence and time to find the right procedure. Inspect the final wording, not only retrieval hits; a correct passage can still be summarized incorrectly. Shadow the workflow on past incidents, then release with source links and a clear unanswered state. Keep the prior index generation so a bad revision import can be rolled back.
Implementation
def release_answer(required_facets, evidence_by_facet, active_passages):
missing = sorted(facet for facet in required_facets
if not evidence_by_facet.get(facet))
if missing:
return {"state": "review", "missing_facets": missing}
for passage_ids in evidence_by_facet.values():
if any(passage_id not in active_passages for passage_id in passage_ids):
return {"state": "review", "reason": "stale-or-restricted-evidence"}
return {"state": "evidence-complete", "facets": sorted(required_facets)}
evidence = {"prerequisite": ["pre-47"], "rollback": ["action-82"],
"verification": ["check-91"]}
assert release_answer(set(evidence), evidence, set(sum(evidence.values(), [])))["state"] == "evidence-complete"
assert release_answer({"rollback", "verification"}, {"rollback": ["action-82"]},
{"action-82"})["state"] == "review"
Performance and operating cost
Checking f facets and p evidence references takes O(f + p) time and O(f) output space, assuming constant-time active-passage membership. Retrieval and human review are larger costs. Evidence completeness is necessary but insufficient: a reviewer must still verify that each cited passage actually supports the sentence and that the steps are ordered safely.
Common Mistakes
- Releasing a rollback answer without its prerequisite.
- Treating a retrieved passage ID as proof of the proposed sentence.
- Mixing instructions from different environments or revisions.
- Hiding a conflict because all required facets have at least one passage.
