Token accuracy can hide missing or shifted entities. Evaluate exact spans, field types and the downstream action that consumes them.
Span-level evaluation and extraction error taxonomy
Score exact matches first
A predicted policy span is correct only when its start, end and field type match the adjudicated gold span under an exact-match contract. A one-character boundary shift can still be useful to a reviewer, but it is not an exact extraction. Report precision, recall and F1 by field with denominators. Do not merge unlabeled or disputed mentions into O without an annotation rule.
Show the disagreement type
Break errors into missed spans, extra spans, boundary shifts, wrong type and repeated-mention confusion. For date extraction, normalization to a calendar value is a separate stage; a correct text span with a wrong normalized date must remain visible. The example counts exact matches and preserves unmatched spans for review.
Keep token scores as diagnostics
Token accuracy can look high when most tokens are O. It helps identify local alignment trouble but does not answer whether the intake system found a policy ID. The correct unit is often a claim-level field: did the system retrieve the right policy identifier anywhere in the note set? State both span and claim-level results when both matter.
Use paired and sliced comparisons
Evaluate two model versions on the same later claims. Break results by note length, depot, note template and OCR quality, with support counts. A decoder change can improve aggregate F1 yet lower recall on handwritten scans. Valid paths are only one part of quality.
Follow the action
A missed incident date might delay a claim; a false policy ID could attach it to the wrong account. Set separate minimum recall or precision for each field and connect those limits to manual review volume. Review thresholds decide which uncertain records need a person.
Implementation
gold = {(7, 12, "POLICY_ID"), (28, 35, "DATE")}
predicted = {(7, 12, "POLICY_ID"), (28, 34, "DATE")}
def exact_span_metrics(expected, actual):
matched = len(expected & actual)
precision = matched / len(actual) if actual else 0.0
recall = matched / len(expected) if expected else 0.0
f1 = 2 * precision * recall / (precision + recall) if precision + recall else 0.0
return {"precision": precision, "recall": recall, "f1": f1,
"missed": expected - actual, "extra": actual - expected}
scores = exact_span_metrics(gold, predicted)
assert scores["precision"] == 0.5
assert scores["recall"] == 0.5
assert (28, 35, "DATE") in scores["missed"]Performance and operating cost
With hashed spans, exact matching is O(G + P) expected time and O(G + P) memory for G gold and P predicted spans. Claim-level matching and adjudication cost more because repeated mentions and normalized values must be reconciled. Keep those costs visible in evaluation planning.
Common Mistakes
- Do not report token accuracy as the only extraction metric.
- Do not silently repair malformed gold labels during scoring.
- Do not count a correctly located but wrongly typed span as exact.
