Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Keyphrase candidates: spans, nesting and document context

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Useful keyphrases are document-grounded spans, not a bag of frequent words or an invented label.

Specify what a phrase represents

A runbook phrase may name a symptom, subsystem or repair step. Decide whether a phrase must occur verbatim in the source, whether abbreviations may be expanded and whether a named identifier belongs in the phrase list. Store the original half-open span and text revision for extractive candidates. “Queue lag” and “queue lag threshold” overlap, but may answer different search needs. Span annotations give reviewers an exact location.

Generate bounded candidates

Start with noun-like spans, domain lexicon terms and short n-grams, then discard boilerplate headers and private identifiers. Preserve sentence and section context so “retry limit” in a warning is not treated as the same instruction as a recommended retry. Candidate rules are allowed to overgenerate; the ranker and review policy decide which phrases are useful. Keep a reason code for each candidate source, and test tokenization on hyphens, punctuation and mixed scripts.

Resolve overlap deliberately

Normalize whitespace and case for duplicate detection without losing source offsets. If both a nested and full phrase remain, give each a purpose: the shorter term may support broad navigation, while the longer term disambiguates a specific failure. Do not concatenate nearby high-scoring words into a phrase that never appeared. Terms generated as taxonomy labels belong in a separate field, clearly marked as derived labels.

Evaluate against retrieval

Phrase precision alone can reward obvious document titles while missing details users search for. Use a reviewed list of target questions, check whether phrases improve discovery and track false links between unrelated incidents. Report phrase coverage, redundancy and stability after document revisions. Ranking and utility covers selection; the runbook project tests the full index.

Implementation

python
def collect_phrase_spans(document, candidate_spans):
    seen = set()
    phrases = []
    for start, end in candidate_spans:
        if not 0 <= start < end <= len(document):
            raise ValueError("invalid source span")
        phrase = document[start:end]
        identity = " ".join(phrase.casefold().split())
        if identity and identity not in seen:
            seen.add(identity)
            phrases.append({"start": start, "end": end, "text": phrase})
    return phrases

runbook = "Inspect queue lag before changing the retry limit."
assert collect_phrase_spans(runbook, [(8, 17), (8, 17)])[0]["text"] == "queue lag"

Performance and operating cost

Validating c spans and normalizing their total character length L takes O(c + L) expected time and O(c + L) space. Generating every n-gram up to a fixed width is linear in document tokens times that width; unconstrained spans grow quadratically. Bound candidate length, and measure downstream reviewer and index cost rather than ranking only phrase count.

Common Mistakes

  • Presenting a generated taxonomy label as a verbatim source phrase.
  • Removing nested phrases without checking their separate search value.
  • Changing normalization after recording offsets.
  • Ranking boilerplate section titles above specific repair terms.

Read next

ai-data
natural-language-processing
Storage details