Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Keyphrase ranking: usefulness, redundancy and drift

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A phrase list should help an agent find the right document. Score that outcome rather than rewarding frequent but empty terms.

Separate salience from utility

A phrase can be prominent in a document and useless in a search result. “Customer issue” may occur often but distinguishes nothing; “callback timeout after 47 retries” may occur once and identify the case. Combine local evidence, corpus rarity, location and domain policy, then inspect results at a fixed phrase budget. Do not fit corpus statistics on held-out documents. Candidate spans determine what the ranker may choose.

Control redundancy and leakage

Deduplicate normalized repeats and compare nested phrases under a policy that values specificity without hiding a useful broad label. Keep document family and tenant boundaries in corpus statistics. A rare customer ID is not a valuable keyphrase; redact or exclude it before ranking. If a phrase list is displayed publicly, ensure every phrase is safe to show and grounded in an accessible source. Redaction is an upstream gate.

Measure with human tasks

Ask reviewers whether a phrase helps identify, distinguish or retrieve a runbook for a realistic question. Also report retrieval at a fixed depth, duplicate phrase rate, private-term leakage and how often phrases survive a document edit. Multiple reviewers may disagree about a useful phrase; keep those judgments rather than flattening them into one hidden label. Compare against document title and TF-IDF baselines on a frozen query set.

Handle revisions

Phrase spans and corpus frequencies can change when the source or index grows. Attach a source revision and extractor version. Recompute affected documents after edits, and withdraw phrases from deleted content. A new ranker should be evaluated on old and newly introduced sections before an alias moves. The project gives the index a reversible release gate.

Implementation

python
def rank_keyphrases(candidates, corpus_document_frequency, corpus_size):
    ranked = []
    for candidate in candidates:
        phrase = candidate["text"].casefold()
        document_frequency = corpus_document_frequency.get(phrase, 0)
        rarity = (corpus_size + 1) / (document_frequency + 1)
        score = candidate["local_count"] * rarity
        ranked.append((score, phrase))
    return sorted(ranked, reverse=True)

phrases = [{"text": "queue lag", "local_count": 2},
           {"text": "incident", "local_count": 3}]
ranked = rank_keyphrases(phrases, {"queue lag": 4, "incident": 90}, 120)
assert ranked[0][1] == "queue lag"

Performance and operating cost

Scoring c candidates is O(c) average dictionary work, followed by O(c log c) sorting and O(c) output space. This baseline is intentionally simple; its score is not a calibrated probability. Corpus-frequency refresh can be expensive across many revisions. Measure retrieval gain, reviewer usefulness and leakage separately before paying for a more complex ranker.

Common Mistakes

  • Treating a high extraction score as proof a phrase helps a user.
  • Learning document frequency from the held-out audit set.
  • Letting a rare customer identifier become the top phrase.
  • Leaving phrases from a deleted source in a public index.

Read next

ai-data
natural-language-processing
Storage details