Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Decode entities and evaluate exact spans, not token accuracy

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A model that tags most words O can look excellent while failing every entity request. Decode spans on the original text and score what downstream code consumes.

Decode against the source message

The model emits labels over tokens or subwords. The consumer needs typed character spans. Walk the ordered token offsets; open a span on B-TYPE, extend it only on I-TYPE of the same type, and close it on O or a new B. Validate that offsets are monotonic and inside the unmodified source string. Store start inclusive and end exclusive. This avoids punctuation guesses when extracting “ZX-47,” an address with a combining mark, or a reference followed by a newline. BIO alignment determines which subword predictions are eligible for decoding.

Count the right errors

Compare predicted and reviewed triples of type, start, and end. Exact match rewards only a correct label with both boundaries correct. Count true positives, false positives and false negatives across the evaluation set before calculating micro precision and recall. Also report per-type and per-language recall. Token accuracy and partial overlap can be useful diagnostics, but neither should be the release gate if the downstream system copies exact substrings into a case record. Boundary errors are operational errors.

Abstention and overlap

Some entity types legitimately overlap: a shipping address may contain a postal code. A flat BIO sequence cannot represent both simultaneously. Either define a single priority type, train separate heads, or use a span model. The choice belongs in the annotation policy. Confidence needs calibration at the extracted-span level if low-confidence entities trigger manual review. The score of one token alone is a poor proxy for a multi-token span. See classification slices and abstention for a related release boundary.

Make failure review reproducible

Keep an immutable evaluation manifest with message identity, language slice, annotation version, tokenizer version and model digest. For each failure, show the original text with gold and predicted offsets; do not log raw personal data into a general dashboard. Separately count malformed transitions, truncated spans, zero-width spans and cross-type overlaps. A changed threshold can alter precision without changing the model, so version the threshold as part of the serving contract.

Implementation

python
def exact_span_counts(gold_by_message, predicted_by_message):
    true_positive = false_positive = false_negative = 0
    message_ids = set(gold_by_message) | set(predicted_by_message)
    for message_id in message_ids:
        gold = set(gold_by_message.get(message_id, ()))
        predicted = set(predicted_by_message.get(message_id, ()))
        true_positive += len(gold & predicted)
        false_positive += len(predicted - gold)
        false_negative += len(gold - predicted)
    precision = true_positive / (true_positive + false_positive) if true_positive + false_positive else 0.0
    recall = true_positive / (true_positive + false_negative) if true_positive + false_negative else 0.0
    return {"precision": precision, "recall": recall,
            "f1": 2 * precision * recall / (precision + recall) if precision + recall else 0.0}

reviewed = {"case-47": {("ORDER", 14, 19), ("POSTCODE", 30, 36)}}
model_output = {"case-47": {("ORDER", 14, 19), ("POSTCODE", 30, 35)}}
assert exact_span_counts(reviewed, model_output)["f1"] == 0.5

Performance and operating cost

With hashable span triples, evaluation takes O(g + p) expected time and O(g + p) space for g gold and p predicted spans. A decoder over t token offsets takes O(t) time. Report input length and throughput per language because token expansion can make one language materially slower than another. Keep sensitive text out of aggregate metric stores; retain restricted samples only for adjudication.

Common Mistakes

  • Using token accuracy as the only extraction metric.
  • Accepting a nearly correct boundary as exact match.
  • Merging adjacent same-type entities without a new B tag.
  • Reporting a single aggregate score when one rare entity type has collapsed.

Read next

Continue the workflow: Project: ship an auditable support-entity extractor.

Continue the workflow: Relation extraction with direction, negation and evidence.

Continue the workflow: OCR field review: confidence, consistency and abstention.

Continue the workflow: Entity linking: candidate generation and the NIL decision.

ai-data
natural-language-processing
Storage details