A token classifier can report high accuracy while losing the exact order number a support agent needs. Define span and subword ownership before training.
BIO labels and subword alignment for entity extraction
The annotation is a span, not a token list
A return request may contain “order AX-47” in a longer sentence. Store the original message, the entity type, and half-open character offsets that select AX-47. Token labels are a derived training view. They depend on the tokenizer version, which means they can change when vocabulary or normalization changes even though the human annotation did not. Keep the character span as the durable record. Span-offset annotation explains the boundary contract; Unicode and tokenization covers why a display character and a code-point offset need careful handling.
BIO transition rules
B-ORDER begins an order identifier, I-ORDER continues it, and O means that the token is outside any tracked span. An I-ORDER after O is malformed under a strict BIO contract. An I-ADDRESS after B-ORDER is malformed too. Choose a policy: reject invalid labels during dataset intake, or repair them at inference with a documented decoder. Silent repair during training hides annotation defects. Adjacent entities of the same type need a fresh B tag. Otherwise two order identifiers can collapse into one prediction.
Subwords and ignored positions
A tokenizer may split AX-47 into several pieces. If the annotation is word-level, assign its label to the first subword and mark the remaining pieces as ignored for loss and metric calculation, or adopt a full-subword labeling policy and keep it consistent everywhere. Special tokens and padding also need ignored targets. The ignored value is a training convention, not an entity class. Record tokenizer digest, label map, alignment policy, maximum length and truncation policy with each trained checkpoint. Recompute derived labels when that bundle changes.
Audit the boundary cases
Create a small fixed set covering punctuation attached to IDs, two adjacent IDs, mixed scripts, whitespace-only messages, truncated tails and spans ending at the last character. Check that every training label can be projected back to the exact original substring. A row that cannot round-trip should be quarantined. This is more useful than a larger set of easy examples. Feed the resulting labels through span decoding and evaluation before using them in the extraction project.
Implementation
def align_first_subword(word_labels, subword_to_word):
aligned = []
previous_word = None
for word_index in subword_to_word:
if word_index is None:
aligned.append(-100)
elif word_index != previous_word:
if not 0 <= word_index < len(word_labels):
raise ValueError("subword points outside the annotated words")
aligned.append(word_labels[word_index])
else:
aligned.append(-100)
previous_word = word_index
return aligned
order_labels = ["O", "B-ORDER", "O"]
piece_owners = [None, 0, 1, 1, 2, None]
assert align_first_subword(order_labels, piece_owners) == [
-100, "O", "B-ORDER", -100, "O", -100
]
Performance and operating cost
Alignment is O(t) time and O(t) output space for t subword positions. Training memory is governed mainly by model activations and sequence length; doubling the maximum length can increase attention work roughly fourfold for full attention. Truncation can erase a labeled span, so report the discarded-span rate and reject or window those rows. At serving time, tokenization and offset mapping must use the same version as training.
Common Mistakes
- Treating a continuation subword as a second entity occurrence.
- Scoring ignored positions as if they were O labels.
- Normalizing text after offsets have been recorded without rebuilding the alignment.
- Allowing an illegal I tag to pass through the training set silently.
Read next
- Entity spans: align annotations to the original text
- Unicode and tokenization: preserve meaning at the text boundary
- Decode entities and evaluate exact spans, not token accuracy
- Project: ship an auditable support-entity extractor
- Text corpus contracts: identity, label timing and annotation rules
Continue the workflow: Decode entities and evaluate exact spans, not token accuracy.
