Build a runbook assistant that returns an authorized, current source span for a direct question and says when the available documents do not answer it.
Project: answer runbook questions with source spans and abstention
Write the response contract
The on-call engineer asks for a retry limit, owner or rollback condition. The response contains answer text, source title, document revision, original offsets and a current-access check. If the source does not state the answer, return a no-answer result and a search link. Do not fabricate a procedure or merge two contradictory versions into one instruction. The user must be able to inspect the evidence.
Prepare retrieval and reader data
Index active runbooks under their revision and permissions. Assemble reviewed questions with exact supporting spans plus unanswerable questions, outdated-version traps and near-miss passages. Group related incidents and copied runbook sections before splitting. Runbook search supplies candidate passages; the span reader validates output against original text.
Gate both stages
Measure retrieval recall at the reader depth, exact-span accuracy on evidence-present questions, unsupported-answer rate, no-answer precision and authorization failures. Tune abstention on validation data; freeze thresholds for the final audit. A reader improvement that relies on feeding it gold passages does not qualify as an end-to-end gain. Answerability calibration determines when to stop.
Operate a changing corpus
Rebuild or update search indexes when runbooks change, invalidate answers tied to retired revisions and preserve rollback. Recheck document permissions for cached results. Sample confident answers for human review and collect cases where the engineer had to search manually. Keep query logs under the incident retention policy; raw questions may contain secrets. A stale answer should become unavailable even if its text still appears in an old cache.
Implementation
def packaged_answer(source_text, source_revision, start, end, authorized):
if not authorized:
return {"state": "abstain", "reason": "access-denied"}
if not source_revision or not 0 <= start < end <= len(source_text):
return {"state": "abstain", "reason": "invalid-source-span"}
return {"state": "answer", "text": source_text[start:end],
"source_revision": source_revision, "start": start, "end": end}
source = "Retry limit: 47 attempts."
assert packaged_answer(source, "runbook-v7", 13, 24, True)["text"] == "47 attempts"
Performance and operating cost
Packaging is O(a) time and space for an a-character answer. End-to-end latency combines index lookup, permission filtering, reader windows and source revalidation. Cache permitted results only with revision and access scope in the key. Track abstention and stale-cache prevention as first-class outcomes; lowering latency by skipping freshness checks can produce a worse operational system.
Common Mistakes
- Returning a snippet from a retired runbook as current guidance.
- Caching by question text alone across users with different permissions.
- Testing the reader only with gold passages.
- Treating a source pointer as proof the selected words answer the actual question.
Read next
- Extractive QA: answer spans, context windows and no-answer cases
- QA answerability: calibrate abstention and evidence quality
- Project: release an incident-runbook search service
- Hybrid text ranking with access filters and reranking
- Document summaries with sentence-level evidence contracts
Continue the workflow: Project: release evidence-checked runbook answers.
Continue the workflow: Project: audit a multi-step incident answer from event to rule.
