Build a reviewed answer path that retrieves authorized passages, marks unsupported claims and abstains when the runbook has no answer.
Project: release evidence-checked runbook answers
Set the answer boundary
An operator asks whether to retry a failed callback. The service retrieves versioned runbook passages under the operator’s permissions and drafts an answer. It must not execute the retry. Every instruction or status claim needs a passage that supports it; if no accessible document answers the question, return an unanswered state. Keep the question, retrieved revision set, generated draft and reviewer decision as separate records.
Create adversarial questions
Include a current runbook, a retired version with a conflicting threshold, a similar service with a different retry limit and an inaccessible incident note. Add questions whose answer is genuinely absent. Reviewers identify supported claims and dangerous commands without seeing the model’s preferred answer first. Claim evidence defines the output; the evaluation rubric defines the release test.
Build the staged pipeline
Retrieve with access filters, assemble passages with IDs and revisions, then generate a short answer. Split it into claims, validate citations and exact operational tokens, and route unsupported or contradictory statements to review. A citation that points to a stale passage should fail even when its text looks plausible. Show the reviewer the passage beside the claim, not merely a confidence number.
Operate and roll back
Shadow-test on recent questions before surfacing drafts to operators. Track wrong-action claims, unsupported facts, unanswered rate, review minutes and retrieval misses. Version index, access policy, prompt and model together. When a runbook changes, invalidate cached answers tied to its old revision. A rollback restores the prior approved bundle while retaining the audit of rejected drafts.
Implementation
def stage_runbook_answer(retrieved, claims, allowed_revisions):
current = {passage["id"] for passage in retrieved
if passage["revision"] in allowed_revisions
and passage["accessible"]}
if not current:
return {"state": "unanswered", "reason": "no-current-evidence"}
for claim in claims:
if claim["passage_id"] not in current or not claim["supported"]:
return {"state": "review", "reason": "unsupported-claim"}
return {"state": "draft-review", "claims": len(claims)}
passages = [{"id": "rb-47", "revision": "r8", "accessible": True}]
claims = [{"passage_id": "rb-47", "supported": True}]
assert stage_runbook_answer(passages, claims, {"r8"})["state"] == "draft-review"
Performance and operating cost
Building the accessible passage set is O(p) time and space for p retrieved passages; checking c claims is O(c) average time. Retrieval and generation dominate machine cost, while review dominates safety cost. Keep the answer short and stop early on missing evidence, because extra prose adds unsupported-claim surface and reviewer work.
Common Mistakes
- Treating an inaccessible passage as acceptable model context.
- Letting an old runbook revision support a current command.
- Skipping review because every sentence has some citation.
- Caching an answer after its supporting runbook is retired.
Read next
- Generated answers: claim evidence and unanswered state
- Generation evaluation: claim support, disagreement and regressions
- Project: answer runbook questions with source spans and abstention
- Hybrid text ranking with access filters and reranking
- QA answerability: calibrate abstention and evidence quality
Continue the workflow: Project: index runbook sections with revision-safe passages.
Continue the workflow: Project: answer long-runbook questions with complete evidence.
Continue the workflow: Project: release a runbook revision without stale answers.
Continue the workflow: Trusted action gates for retrieval-assisted answers.
