An answer reader should return the exact source span or abstain. Long documents, repeated phrases and missing evidence make that contract harder than selecting two token positions.
Extractive QA: answer spans, context windows and no-answer cases
Define the answer object
Store question, source document revision, supporting passage, half-open answer offsets and an answerability label. An extractive answer must select bytes or characters from the original representation agreed by the corpus policy; token start and end are a temporary model view. A question can have several correct spans or no answer in the supplied context. Do not mark a plausible outside fact as correct when the retrieved document does not say it. Span annotation supplies the source coordinate contract.
Window long context deliberately
A reader has a bounded token budget. Split a long document into overlapping windows and keep each window’s mapping to original offsets. During training, a window without the gold answer is a no-answer window even if another window contains it. At inference, compare candidates across windows under a calibrated rule and deduplicate identical spans. A narrow stride can lose an answer at a boundary; a large overlap adds latency and correlated false positives. Save the window policy with the model bundle.
Separate retrieval from reading
The reader cannot recover an answer absent from retrieved passages. Measure retrieval recall at candidate depth, answerable-case span F1 and no-answer precision separately. A strong reader score on gold passages can conceal a weak search stage. Use runbook retrieval to provide permitted candidate passages and access filtering before any reader sees private text.
Audit the hard questions
Include repeated numbers, two dates, negated instructions, table-like text, outdated revisions and questions requiring several passages. Exact match is strict and useful for identifiers; token overlap can diagnose near misses but cannot certify an amount or date. A source-span answer should travel with document revision and confidence. Answerability calibration decides when to return it, while the project tests the whole path.
Implementation
def answer_from_source(source_text, start, end, expected=None):
if not 0 <= start < end <= len(source_text):
raise ValueError("answer span outside source")
answer = source_text[start:end]
if expected is not None and answer != expected:
raise ValueError("answer no longer matches source revision")
return {"answer": answer, "start": start, "end": end}
runbook = "The callback retry limit is 47 attempts."
assert answer_from_source(runbook, 28, 39, "47 attempts")["answer"] == "47 attempts"
Performance and operating cost
Span validation is O(a) time to copy an a-character answer and O(a) output space. With w context windows, reader inference costs roughly w model passes; overlap raises both w and duplicate-candidate work. Retrieval and access checks add latency before reading begins. Measure p95 per question length and document length, plus the fraction of answerable questions whose evidence never reached the reader.
Common Mistakes
- Scoring the reader only on passages guaranteed to contain an answer.
- Assigning an answer to a window that does not contain its gold span.
- Returning a plausible fact from model memory when the source is silent.
- Dropping source revision or original offsets from the answer object.
Read next
- QA answerability: calibrate abstention and evidence quality
- Project: answer runbook questions with source spans and abstention
- Entity spans: align annotations to the original text
- Project: release an incident-runbook search service
- Hybrid text ranking with access filters and reranking
Continue the workflow: QA answerability: calibrate abstention and evidence quality.
