Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Sentence boundaries in technical text: candidates and abbreviations

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Find sentence boundaries without breaking decimals, service versions or domain abbreviations, and retain every character offset.

Boundary detection is a separate decision

Tokenization can split words while still leaving the sentence boundary uncertain. In incident notes, a period may end a sentence, close an abbreviation or separate digits in a version. “svc. gateway-west failed” should not automatically become two claims. Propose punctuation candidates, then decide each one from context and a versioned domain policy. Keep character offsets into the original note so later entities and claims can be traced. Unicode tokenization explains why this is more than splitting on spaces.

Use domain evidence carefully

A short list of approved abbreviations can prevent obvious false splits. It should include the local meaning and revision because “Dr.” in a support note may differ from a release label. A pattern for decimals helps only when digits surround the period; it does not resolve every product identifier. URLs, file paths, dotted service names and code blocks need their own protected spans. Avoid a blanket rule that never splits after an abbreviation, because an abbreviation can also occur at the end of a sentence. Abbreviation scope provides the registry model.

Preserve text, not reconstructed guesses

Store start and end offsets for each sentence. Joining trimmed fragments back with a single space changes the source and can invalidate evidence spans. If a boundary is ambiguous, keep the candidate and mark it for review rather than inventing punctuation. The example below offers a narrow plain-text baseline: it recognizes punctuation followed by whitespace, suppresses known abbreviations and leaves the final segment intact. It is not an HTML or code-aware parser.

Evaluate downstream damage

Count boundary precision and recall, but also test whether a wrong split changes entity extraction, negation scope or claim attribution. Include logs, chat fragments, decimal measurements, ellipses and multilingual punctuation. Measure errors by document family; a high overall score can hide failures on the operational notes that matter most. The regression lesson explains offset checks, and the project gates release.

Implementation

python
import re

def conservative_sentences(note, abbreviations):
    spans = []
    start = 0
    for marker in re.finditer(r"[.!?](?=\s|$)", note):
        prefix = note[start:marker.end()]
        last_token = prefix.rstrip().split()[-1] if prefix.strip() else ""
        if last_token in abbreviations:
            continue
        end = marker.end()
        spans.append((start, end))
        start = end
        while start < len(note) and note[start].isspace():
            start += 1
    if start < len(note):
        spans.append((start, len(note)))
    return spans

note = "Escalate to ops. svc. gateway-west recovered at 03:47."
spans = conservative_sentences(note, {"svc."})
assert [note[a:b] for a, b in spans] == [
    "Escalate to ops.", "svc. gateway-west recovered at 03:47."]

Performance and operating cost

For n characters and b candidate marks, the regex scan is O(n), while slicing prefixes can make this simple demonstration O(n²) in a long note. Store a rolling token boundary or use a trained segmenter for large inputs. Output spans use O(b) space. The code intentionally leaves leading spaces outside spans; a production offset contract must state how those gaps are represented.

Common Mistakes

  • Splitting on every period.
  • Treating every abbreviation as permanently nonterminal.
  • Discarding source offsets after trimming whitespace.
  • Testing only clean prose while shipping on logs and incident notes.

Read next

ai-data
natural-language-processing
Storage details