Find sentence boundaries without breaking decimals, service versions or domain abbreviations, and retain every character offset.
Sentence boundaries in technical text: candidates and abbreviations
Boundary detection is a separate decision
Tokenization can split words while still leaving the sentence boundary uncertain. In incident notes, a period may end a sentence, close an abbreviation or separate digits in a version. “svc. gateway-west failed” should not automatically become two claims. Propose punctuation candidates, then decide each one from context and a versioned domain policy. Keep character offsets into the original note so later entities and claims can be traced. Unicode tokenization explains why this is more than splitting on spaces.
Use domain evidence carefully
A short list of approved abbreviations can prevent obvious false splits. It should include the local meaning and revision because “Dr.” in a support note may differ from a release label. A pattern for decimals helps only when digits surround the period; it does not resolve every product identifier. URLs, file paths, dotted service names and code blocks need their own protected spans. Avoid a blanket rule that never splits after an abbreviation, because an abbreviation can also occur at the end of a sentence. Abbreviation scope provides the registry model.
Preserve text, not reconstructed guesses
Store start and end offsets for each sentence. Joining trimmed fragments back with a single space changes the source and can invalidate evidence spans. If a boundary is ambiguous, keep the candidate and mark it for review rather than inventing punctuation. The example below offers a narrow plain-text baseline: it recognizes punctuation followed by whitespace, suppresses known abbreviations and leaves the final segment intact. It is not an HTML or code-aware parser.
Evaluate downstream damage
Count boundary precision and recall, but also test whether a wrong split changes entity extraction, negation scope or claim attribution. Include logs, chat fragments, decimal measurements, ellipses and multilingual punctuation. Measure errors by document family; a high overall score can hide failures on the operational notes that matter most. The regression lesson explains offset checks, and the project gates release.
Implementation
import re
def conservative_sentences(note, abbreviations):
spans = []
start = 0
for marker in re.finditer(r"[.!?](?=\s|$)", note):
prefix = note[start:marker.end()]
last_token = prefix.rstrip().split()[-1] if prefix.strip() else ""
if last_token in abbreviations:
continue
end = marker.end()
spans.append((start, end))
start = end
while start < len(note) and note[start].isspace():
start += 1
if start < len(note):
spans.append((start, len(note)))
return spans
note = "Escalate to ops. svc. gateway-west recovered at 03:47."
spans = conservative_sentences(note, {"svc."})
assert [note[a:b] for a, b in spans] == [
"Escalate to ops.", "svc. gateway-west recovered at 03:47."]
Performance and operating cost
For n characters and b candidate marks, the regex scan is O(n), while slicing prefixes can make this simple demonstration O(n²) in a long note. Store a rolling token boundary or use a trained segmenter for large inputs. Output spans use O(b) space. The code intentionally leaves leading spaces outside spans; a production offset contract must state how those gaps are represented.
Common Mistakes
- Splitting on every period.
- Treating every abbreviation as permanently nonterminal.
- Discarding source offsets after trimming whitespace.
- Testing only clean prose while shipping on logs and incident notes.
Read next
- Sentence segmentation: offset contracts and regression slices
- Project: segment incident notes without losing evidence spans
- Unicode and tokenization: preserve meaning at the text boundary
- Abbreviations: definition scope, collisions and unknown forms
- Predicate arguments: actors, targets and negation scope
