An abstention score must show how often the system answers and how often those answers are wrong. A model that sends every case to review has no answered-case errors but provides no useful coverage. Score by operational slice rather than collapsing low-risk and high-risk cases into a single average. Freeze expected labels before evaluating a candidate prompt, and keep the examples out of the prompt itself. Here, the label review means the system withheld a final decision.
Code lab: score answered cases and abstentions
Decision in practice
A claims team tests five held-out cases. Two routine cases receive correct final labels. Two high-impact cases include one wrong approval and one deliberate review. A third routine case is also withheld because its receipt date is missing. A single accuracy percentage would conceal the unsupported high-impact approval. The release conversation needs answered counts, wrong answered counts, and withheld counts for each slice, with the critical error called out separately.
cases = [
{"id": "CL-301", "slice": "routine", "expected": "approve", "actual": "approve"},
{"id": "CL-302", "slice": "routine", "expected": "deny", "actual": "deny"},
{"id": "CL-303", "slice": "routine", "expected": "deny", "actual": "review"},
{"id": "CL-304", "slice": "high-impact", "expected": "deny", "actual": "approve"},
{"id": "CL-305", "slice": "high-impact", "expected": "deny", "actual": "review"},
]
for case_slice in ("routine", "high-impact"):
selected = [case for case in cases if case["slice"] == case_slice]
answered = [case for case in selected if case["actual"] != "review"]
wrong = [case for case in answered if case["actual"] != case["expected"]]
print(f"{case_slice}: answered={len(answered)}/{len(selected)} wrong={len(wrong)}")
Expected output: routine: answered=2/3 wrong=0; high-impact: answered=1/2 wrong=1
Performance and operating cost
The simple implementation scans the N cases once per slice, so S slices cost O(SN) time and O(N) temporary space in the largest pass. Grouping cases in one pass can reduce time to O(N + S) when the dataset grows. The computational cost is minor; labeling and adjudicating cases usually dominate. Report raw counts alongside rates because a zero in a tiny slice is weak evidence. An expected review label needs a separate policy for whether withholding was correct; this sample measures coverage and wrong final decisions only.
Common Mistakes
- Do not report an error rate without the answered share.
- Do not combine high-impact and routine errors into one release threshold.
- Do not call a withheld case correct merely because it avoided a wrong answer.
