A token label is meaningful only when the text offsets, tokenization version and span rules agree with the annotation contract.
Sequence-label contracts and token alignment
Define a span in the source text
A claims-intake model extracts policy IDs and incident dates from adjuster notes. Store each gold field as character offsets and a type, with a rule for inclusive start and exclusive end. Word tokens are a training representation, not the source of truth. The string “Policy ZX-47” might be split into several model pieces, but the requested output remains one policy span. A change in Unicode normalization can shift offsets; version that step with the corpus.
Pick a label scheme deliberately
A BIO scheme marks the beginning and inside of each entity, with O outside. It assumes flat, nonoverlapping spans; a policy ID inside a larger address field may require another representation. Declare whether adjacent same-type spans are separate and whether punctuation belongs to the entity. The tagger should reject records that cannot be expressed under the chosen scheme rather than silently dropping an annotation.
Align subwords without multiplying evidence
A tokenizer may divide “ZX-47” into multiple pieces. One defensible training rule labels only the first piece and excludes the remaining pieces from loss; another propagates adjusted labels across pieces. Choose one rule and preserve a map from each piece to source offsets. Scoring must reconstruct original spans, never count three subword decisions as three separate policy IDs.
Split by claim and event time
Notes from one claim often repeat the same identifiers. Put all notes for a claim in one split and reserve later claims for the final test. A random note split lets the model memorize IDs or templates. Group and time validation provides the boundary.
Audit annotation disagreement
Some dates are ambiguous: “next Friday” may be a mention without a resolvable calendar date. Keep span detection separate from normalization, adjudicate disagreements and retain an unknown state. Do not force every text mention into a date value. A later span evaluation needs these rules.
Implementation
note = "Policy ZX-47 was renewed on 18 June."
gold_spans = [(7, 12, "POLICY_ID"), (28, 35, "DATE")]
def validate_spans(text, spans):
previous_end = 0
for start, end, field_type in sorted(spans):
if not (0 <= previous_end <= start < end <= len(text)):
raise ValueError("invalid or overlapping span")
if not field_type or not text[start:end].strip():
raise ValueError("empty field")
previous_end = end
return [text[start:end] for start, end, _ in spans]
assert validate_spans(note, gold_spans) == ["ZX-47", "18 June"]Performance and operating cost
Validating S sorted spans costs O(S log S) for sorting and O(S) thereafter; storing offsets costs O(S). Tokenization and encoder inference add length-dependent cost. More detailed annotation contracts raise labeling effort but prevent hidden disagreements from being counted as model errors.
Common Mistakes
- Do not score raw subword labels as if they were original-text spans.
- Do not silently discard overlapping annotations.
- Do not random-split repeated notes from one claim.
