Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Sentence segmentation: offset contracts and regression slices

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Validate sentence spans against the exact source revision and measure boundary errors where they alter later NLP decisions.

Define the offset convention

A segmentation result is an ordered set of source intervals, not a list of rewritten strings. State whether intervals include trailing punctuation and surrounding whitespace, and whether indexing uses Unicode code points or bytes. Never mix conventions across OCR, browser text and Python. Store source revision and normalization version next to the intervals. Source-offset mapping handles transformations that change string length.

Keep ambiguous candidates visible

A line such as “Contact svc. gateway-west. Restart pending.” has one abbreviation and one clear boundary. A model can score both, but a low-margin decision should remain reviewable. Do not attach negation, service identity or action frames to a span whose boundary is unstable without checking downstream effect. A false merge can make one sentence’s warning appear to govern the next instruction; a false split can separate a condition from its action.

Build a diverse regression set

Freeze notes from several writing styles: terse on-call fragments, copied stack traces, version strings, decimal durations, ellipses, multilingual punctuation and names ending in a period. Group related notes by incident so near duplicates do not leak between training and test. Evaluate exact boundary positions and downstream entity and claim outputs. Family-aware splits help prevent an easy but misleading score.

Check intervals mechanically

Each non-whitespace character should belong to one and only one sentence under the chosen convention. Gaps and overlaps are errors even if the displayed sentence strings look plausible. The small check below validates ordered nonoverlap and source coverage after ignoring whitespace gaps; semantic correctness still requires annotated boundaries. The project uses both mechanical and task-level checks.

Implementation

python
def verify_sentence_spans(note, spans):
    cursor = 0
    for start, end in spans:
        if not 0 <= start < end <= len(note):
            return False
        if note[cursor:start].strip() or start < cursor:
            return False
        cursor = end
    return not note[cursor:].strip()

note = "Alarm cleared.  Follow-up pending."
spans = [(0, 14), (16, len(note))]
assert verify_sentence_spans(note, spans)
assert not verify_sentence_spans(note, [(0, 14)])
assert not verify_sentence_spans(note, [(0, 14), (12, len(note))])

Performance and operating cost

For s spans, index checks are O(s); checking gap substrings across the whole note costs O(n) total for n characters when spans are ordered. Space is O(1) beyond the input. This verifies coverage, not whether a period is the correct linguistic boundary. Use gold offsets and downstream task metrics for that judgment.

Common Mistakes

  • Calling contiguous spans accurate without gold boundaries.
  • Mixing UTF-8 byte positions with Unicode character positions.
  • Allowing a skipped non-whitespace fragment to disappear.
  • Reporting one aggregate score without noisy-note slices.

Read next

ai-data
natural-language-processing
Storage details