Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Code lab: score answered cases and abstentions

Last updated: 2 Oct 20269 min read
tutorial
AdvancedBy AITrove Editorial

An abstention score must show how often the system answers and how often those answers are wrong. A model that sends every case to review has no answered-case errors but provides no useful coverage. Score by operational slice rather than collapsing low-risk and high-risk cases into a single average. Freeze expected labels before evaluating a candidate prompt, and keep the examples out of the prompt itself. Here, the label review means the system withheld a final decision.

Decision in practice

A claims team tests five held-out cases. Two routine cases receive correct final labels. Two high-impact cases include one wrong approval and one deliberate review. A third routine case is also withheld because its receipt date is missing. A single accuracy percentage would conceal the unsupported high-impact approval. The release conversation needs answered counts, wrong answered counts, and withheld counts for each slice, with the critical error called out separately.

python
cases = [
    {"id": "CL-301", "slice": "routine", "expected": "approve", "actual": "approve"},
    {"id": "CL-302", "slice": "routine", "expected": "deny", "actual": "deny"},
    {"id": "CL-303", "slice": "routine", "expected": "deny", "actual": "review"},
    {"id": "CL-304", "slice": "high-impact", "expected": "deny", "actual": "approve"},
    {"id": "CL-305", "slice": "high-impact", "expected": "deny", "actual": "review"},
]

for case_slice in ("routine", "high-impact"):
    selected = [case for case in cases if case["slice"] == case_slice]
    answered = [case for case in selected if case["actual"] != "review"]
    wrong = [case for case in answered if case["actual"] != case["expected"]]
    print(f"{case_slice}: answered={len(answered)}/{len(selected)} wrong={len(wrong)}")

Expected output: routine: answered=2/3 wrong=0; high-impact: answered=1/2 wrong=1

Performance and operating cost

The simple implementation scans the N cases once per slice, so S slices cost O(SN) time and O(N) temporary space in the largest pass. Grouping cases in one pass can reduce time to O(N + S) when the dataset grows. The computational cost is minor; labeling and adjudicating cases usually dominate. Report raw counts alongside rates because a zero in a tiny slice is weak evidence. An expected review label needs a separate policy for whether withholding was correct; this sample measures coverage and wrong final decisions only.

Common Mistakes

  • Do not report an error rate without the answered share.
  • Do not combine high-impact and routine errors into one release threshold.
  • Do not call a withheld case correct merely because it avoided a wrong answer.

Connected lessons

prompt engineering
evaluation
python
Storage details