A readable answer can still assert a fact absent from the retrieved material. Keep support and refusal as explicit outcomes.
Generated answers: claim evidence and unanswered state
Decompose the output into claims
For a runbook answer, identify each operational claim: current status, cause, action and condition. A claim may be supported, contradicted or not addressed by accessible passages. Store passage IDs and revisions for the supporting span. Do not treat mere word overlap as evidence. If the retrieved material does not answer a question, return an unanswered state and a useful next step. Entailment gives a related premise-hypothesis test.
Constrain generation before review
Pass only passages the requester can access. Keep exact commands, identifiers and version requirements separate from prose so they can be checked mechanically. Ask the model to produce a concise answer with claim-to-passage links, but validate those links after generation. A cited passage might exist yet not support the sentence attached to it. Retrieval filters must run before model context is assembled.
Gate risky assertions
Block unsupported instructions that alter production state, financial outcomes or customer commitments. For lower-risk content, show a draft with marked unsupported sentences and let an authorized reviewer decide. A low-confidence answer is not repaired by adding a generic warning while preserving the false claim. Prefer no answer when the evidence is absent or stale. Regression evaluation measures these cases.
Separate fluency from utility
Compare claim support, answerability, task completion and review time. A fluent paragraph may be worse than a brief “not documented” response when the runbook is missing. Include conflicting source revisions and retrieved distractors in the audit. The runbook release project stages answers with explicit supported and unresolved states.
Implementation
def claim_release_state(claims, accessible_passage_ids):
if not claims:
return "unanswered"
for claim in claims:
if claim["risk"] == "action" and claim["support"] != "supported":
return "blocked"
if claim["support"] != "supported":
return "review"
if claim["passage_id"] not in accessible_passage_ids:
return "blocked"
return "draft-review"
claims = [{"risk": "info", "support": "supported", "passage_id": "rb-47"}]
assert claim_release_state(claims, {"rb-47"}) == "draft-review"
assert claim_release_state([], {"rb-47"}) == "unanswered"
Performance and operating cost
The gate scans c claims in O(c) time and uses O(1) extra space with a set of accessible passage IDs. Claim extraction, retrieval and human verification cost more. More generated sentences create more claims to inspect, so concise output can lower both inference and review expense without sacrificing the answer.
Common Mistakes
- Treating a linked passage as proof it supports a claim.
- Answering from inaccessible or retired runbook material.
- Appending a warning to an unsupported production command instead of blocking it.
- Reporting fluency while ignoring unanswered and false-action cases.
Read next
- Generation evaluation: claim support, disagreement and regressions
- Project: release evidence-checked runbook answers
- Textual entailment: premise, hypothesis and evidence boundaries
- Hybrid text ranking with access filters and reranking
- Verify summary claims, corrections and human-review triggers
Continue the workflow: Revision impact: invalidate derived claims, indexes and answers.
Continue the workflow: Retrieved text is evidence, not an instruction source.
