A phrase list should help an agent find the right document. Score that outcome rather than rewarding frequent but empty terms.
Keyphrase ranking: usefulness, redundancy and drift
Separate salience from utility
A phrase can be prominent in a document and useless in a search result. “Customer issue” may occur often but distinguishes nothing; “callback timeout after 47 retries” may occur once and identify the case. Combine local evidence, corpus rarity, location and domain policy, then inspect results at a fixed phrase budget. Do not fit corpus statistics on held-out documents. Candidate spans determine what the ranker may choose.
Control redundancy and leakage
Deduplicate normalized repeats and compare nested phrases under a policy that values specificity without hiding a useful broad label. Keep document family and tenant boundaries in corpus statistics. A rare customer ID is not a valuable keyphrase; redact or exclude it before ranking. If a phrase list is displayed publicly, ensure every phrase is safe to show and grounded in an accessible source. Redaction is an upstream gate.
Measure with human tasks
Ask reviewers whether a phrase helps identify, distinguish or retrieve a runbook for a realistic question. Also report retrieval at a fixed depth, duplicate phrase rate, private-term leakage and how often phrases survive a document edit. Multiple reviewers may disagree about a useful phrase; keep those judgments rather than flattening them into one hidden label. Compare against document title and TF-IDF baselines on a frozen query set.
Handle revisions
Phrase spans and corpus frequencies can change when the source or index grows. Attach a source revision and extractor version. Recompute affected documents after edits, and withdraw phrases from deleted content. A new ranker should be evaluated on old and newly introduced sections before an alias moves. The project gives the index a reversible release gate.
Implementation
def rank_keyphrases(candidates, corpus_document_frequency, corpus_size):
ranked = []
for candidate in candidates:
phrase = candidate["text"].casefold()
document_frequency = corpus_document_frequency.get(phrase, 0)
rarity = (corpus_size + 1) / (document_frequency + 1)
score = candidate["local_count"] * rarity
ranked.append((score, phrase))
return sorted(ranked, reverse=True)
phrases = [{"text": "queue lag", "local_count": 2},
{"text": "incident", "local_count": 3}]
ranked = rank_keyphrases(phrases, {"queue lag": 4, "incident": 90}, 120)
assert ranked[0][1] == "queue lag"
Performance and operating cost
Scoring c candidates is O(c) average dictionary work, followed by O(c log c) sorting and O(c) output space. This baseline is intentionally simple; its score is not a calibrated probability. Corpus-frequency refresh can be expensive across many revisions. Measure retrieval gain, reviewer usefulness and leakage separately before paying for a more complex ranker.
Common Mistakes
- Treating a high extraction score as proof a phrase helps a user.
- Learning document frequency from the held-out audit set.
- Letting a rare customer identifier become the top phrase.
- Leaving phrases from a deleted source in a public index.
