Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Entity spans: align annotations to the original text

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

An extraction label names a half-open character span in a specific text revision, with an explicit policy for overlap and normalization.

Store coordinates with identity

A ticket might mention an invoice ID, product and date. Store start and end character offsets against the exact raw text revision that annotators saw. If a later normalization collapses spaces, old offsets may point to different characters. Save the selected substring and a text hash beside offsets to catch drift. Corpus identity] makes that pairing stable.

Define overlap semantics

A product code may appear inside a longer order reference. Decide whether overlapping spans are allowed and whether nested labels are meaningful. If a model or library supports only non-overlapping entities, do not quietly drop valid nested annotations; change the representation or define a priority policy. State whether offsets count Unicode code points, bytes or another unit at every API boundary.

Align without guessing

Token-based models need character spans mapped to token spans. A strict mapping may reject an annotation cutting through a token. A permissive expansion can change the labeled text. Count rejected alignments and inspect them before training. Tokenization changes] can alter these counts even when raw documents are unchanged.

Evaluate exact and partial errors

Exact span match is strict; a prediction missing one character can be wrong even when its label looks right. Report exact boundary errors, wrong labels and missed spans separately. Add fixtures for repeated identical IDs in one ticket, punctuation next to an ID and a corrected ticket revision. A raw substring check prevents stale offsets from passing silently.

Implementation

python
def validate_span(ticket_text, start_char, end_char, selected_text):
    if not (0 <= start_char < end_char <= len(ticket_text)):
        raise ValueError("span outside ticket")
    observed = ticket_text[start_char:end_char]
    if observed != selected_text:
        raise ValueError("annotation offsets no longer match text")
    return {"start": start_char, "end": end_char, "text": observed}

Performance and operating cost

One span validation is O(L) for substring length L; validating S spans costs the total selected length. Token alignment adds tokenizer work over document length and storage for offset maps.

Common Mistakes

  • Do not reuse offsets after changing text normalization.
  • Do not silently expand misaligned spans.
  • Do not assume repeated strings identify the intended occurrence.

Read next

ai-data
natural-language-processing
Storage details