Skip to content
AITroveRead. Build. Understand.
Make this comfortable

QA answerability: calibrate abstention and evidence quality

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A high-scoring span may answer the wrong question. Compare it with a no-answer option and the quality of the retrieved evidence before replying.

Distinguish three failures

The collection may lack the answer, retrieval may miss an existing passage, or the reader may misread a good passage. These failures require different fixes. Label answerability relative to a frozen source snapshot, and separately label whether the retrieved candidate set contains a supporting passage. A no-answer response is valid when evidence is absent; a missed retrieval is a pipeline defect. The span contract defines what a positive answer must point to.

Calibrate the decision

Fit a threshold on validation questions that includes answerable and unanswerable cases. Compare the best span score with a no-answer score or other calibrated evidence feature, according to the reader design. Do not interpret a raw model logit as a probability. Report coverage, exact-answer precision, unsupported-answer rate and correctly abstained cases at several thresholds. If the cost of a wrong operational instruction is high, require stronger evidence or a human handoff.

Check the source itself

A retrieved passage can mention the right terms yet not answer the question. Verify that the selected span is inside an authorized, current source revision and that the surrounding sentence supports the requested relation. “Retry limit is 47” does not answer “how many retries have already happened?” Consider contradictory passages and later corrections. The evidence contract shows why a citation-like pointer does not by itself prove entailment.

Slice the error budget

Evaluate rare identifiers, ambiguous questions, multi-hop questions, obsolete runbooks and mixed-language phrasing separately. Keep questions from the same incident in one split. Review a sample of confident answers and abstentions after release to avoid learning only from complaints. If source access changes, invalidate cached answers or recheck authorization. The runbook QA project applies these checks to a support workflow.

Implementation

python
def answer_decision(best_span_score, no_answer_score, evidence_current,
                    evidence_authorized, margin=0.17):
    if not evidence_current or not evidence_authorized:
        return "abstain", "invalid-evidence"
    if best_span_score - no_answer_score < margin:
        return "abstain", "insufficient-margin"
    return "answer", "reviewed-threshold-passed"

assert answer_decision(0.69, 0.61, True, True) == (
    "abstain", "insufficient-margin"
)

Performance and operating cost

The decision gate is O(1) time and space. Calibration requires reviewed questions and repeated scoring across candidate thresholds; that offline cost is small relative to labeling hard no-answer cases. At serving time, authorization and freshness checks may add lookup latency but cannot be skipped. Report wrong-answer cost, abstention volume and manual follow-up time with model throughput.

Common Mistakes

  • Calling a raw score a calibrated probability.
  • Treating retrieval failure as a valid unanswerable question.
  • Answering from a stale or unauthorized passage.
  • Reporting accuracy on answered questions without coverage and unsupported-answer rate.

Read next

ai-data
natural-language-processing
Storage details