Validate sentence spans against the exact source revision and measure boundary errors where they alter later NLP decisions.
Sentence segmentation: offset contracts and regression slices
Define the offset convention
A segmentation result is an ordered set of source intervals, not a list of rewritten strings. State whether intervals include trailing punctuation and surrounding whitespace, and whether indexing uses Unicode code points or bytes. Never mix conventions across OCR, browser text and Python. Store source revision and normalization version next to the intervals. Source-offset mapping handles transformations that change string length.
Keep ambiguous candidates visible
A line such as “Contact svc. gateway-west. Restart pending.” has one abbreviation and one clear boundary. A model can score both, but a low-margin decision should remain reviewable. Do not attach negation, service identity or action frames to a span whose boundary is unstable without checking downstream effect. A false merge can make one sentence’s warning appear to govern the next instruction; a false split can separate a condition from its action.
Build a diverse regression set
Freeze notes from several writing styles: terse on-call fragments, copied stack traces, version strings, decimal durations, ellipses, multilingual punctuation and names ending in a period. Group related notes by incident so near duplicates do not leak between training and test. Evaluate exact boundary positions and downstream entity and claim outputs. Family-aware splits help prevent an easy but misleading score.
Check intervals mechanically
Each non-whitespace character should belong to one and only one sentence under the chosen convention. Gaps and overlaps are errors even if the displayed sentence strings look plausible. The small check below validates ordered nonoverlap and source coverage after ignoring whitespace gaps; semantic correctness still requires annotated boundaries. The project uses both mechanical and task-level checks.
Implementation
def verify_sentence_spans(note, spans):
cursor = 0
for start, end in spans:
if not 0 <= start < end <= len(note):
return False
if note[cursor:start].strip() or start < cursor:
return False
cursor = end
return not note[cursor:].strip()
note = "Alarm cleared. Follow-up pending."
spans = [(0, 14), (16, len(note))]
assert verify_sentence_spans(note, spans)
assert not verify_sentence_spans(note, [(0, 14)])
assert not verify_sentence_spans(note, [(0, 14), (12, len(note))])
Performance and operating cost
For s spans, index checks are O(s); checking gap substrings across the whole note costs O(n) total for n characters when spans are ordered. Space is O(1) beyond the input. This verifies coverage, not whether a period is the correct linguistic boundary. Use gold offsets and downstream task metrics for that judgment.
Common Mistakes
- Calling contiguous spans accurate without gold boundaries.
- Mixing UTF-8 byte positions with Unicode character positions.
- Allowing a skipped non-whitespace fragment to disappear.
- Reporting one aggregate score without noisy-note slices.
Read next
- Sentence boundaries in technical text: candidates and abbreviations
- Project: segment incident notes without losing evidence spans
- Unicode normalization: search keys without losing source offsets
- Family-aware text splits and evaluation contamination
- Text classification evaluation: inspect slices and allow abstention
