An abstention policy defines the evidence and validation conditions under which the system may answer, and when it must return review or ask for missing data. Do not treat a model's self-reported confidence number as a calibrated probability without testing it. Build a labeled set that includes clean, conflicting, and insufficient cases. Measure the error rate among answered cases together with the share withheld. Raising the bar may reduce wrong decisions but increase human queue load; choose the tradeoff for the consequence of the task.
Selective answers: measure when to abstain
Decision in practice
A claims assistant can approve, deny, or request review. The release rule requires a current policy clause, matching account, and valid receipt date before either final decision. On 68 held-out cases, it withholds 11 with missing or conflicting facts. Reviewers find one unsupported approval among the 57 answered cases. The team keeps that raw count and investigates it; it does not say the system is safe because most outputs looked cautious. A stricter rule may send more cases to review, so staffing and response time are part of the release decision.
Answer only if: current clause + account match + valid receipt date.
Otherwise: review with missing-field reason.
Report: answered_count, wrong_answer_count, withheld_count.
Gate: critical unsupported approvals must be zero on the release set.Performance and operating cost
Computing rule gates is cheap relative to generation, but withheld cases shift cost to people. Track both coverage and risk: a system that abstains on every case has no wrong answers but does not perform the job. Use separate thresholds for low- and high-consequence actions. Recalibrate after policy or model changes, because a threshold tuned on old cases may no longer select the right set. Do not infer uncertainty from response length or politeness.
Common Mistakes
- Do not call model self-confidence a calibrated probability without testing.
- Do not report error rate without the withheld share.
- Do not force a final label when required evidence is absent.
Connected lessons
- Prompt Engineering
- Production prompt engineering
- Output contracts: parse a result and preserve an explicit unknown state
- Human handoff: preserve evidence and the reason for uncertainty
- Evaluation leakage: keep the release test independent
- Project: measure a retrieval-backed answer gate
- Prompt evidence and output decisions
Try the executable check: Code lab: score answered cases and abstentions.
