A retrained topic model can swap topic numbers, split a theme or merge two themes. Compare reviewed content and assignments before announcing a trend.
Topic drift: align changing clusters before reporting trends
Do not compare raw IDs
Topic 4 this month has no guaranteed relationship with Topic 4 last month. Keep a frozen anchor set of tickets and score both models on it. Compare representative examples, characteristic terms and assignment overlap for candidate alignments. A high overlap is evidence for continuity, not proof: two broad topics may both cover a common template while diverging on rare incidents. Record split, merge and new-theme events explicitly.
Separate real movement from pipeline changes
A rise in “refund delay” may reflect more complaints, a new classification policy, a changed support channel or a tokenizer update. Preserve model, preprocessing and corpus manifests. Estimate trend on a fixed labeled definition when possible; if not, state that the series has a break. Slice by language and product release because a newly added channel can shift the entire corpus mix. Topic discovery explains why exploratory clusters need reviewed names.
Review the alignment
Match candidate old-new topics using overlap on the same anchor tickets and a reviewer’s judgment of the central issue. Do not force a one-to-one map: allow many-to-one merges and one-to-many splits. Keep low-confidence alignments in an “unresolved” bucket until more examples are reviewed. Reuse the same annotated anchor set across versions only while its cases remain representative. Refresh it with a logged policy and keep the older benchmark for regression comparison.
Publish cautious metrics
Report volume, share of eligible tickets, reviewer agreement, topic coverage and the fraction left unresolved. Attach a series version to every plotted point. A business-facing statement should distinguish “more tickets arrived” from “the model reassigned tickets.” Give reviewers the top changed examples, not only a similarity score. The support-theme project uses these gates before a weekly dashboard changes.
Implementation
def overlap_matrix(old_assignments, new_assignments):
if set(old_assignments) != set(new_assignments):
raise ValueError("compare models on the same anchor tickets")
overlap = {}
for ticket_id, old_topic in old_assignments.items():
pair = (old_topic, new_assignments[ticket_id])
overlap[pair] = overlap.get(pair, 0) + 1
return overlap
previous = {"case-47": "retry", "case-82": "retry", "case-91": "refund"}
current = {"case-47": "callback", "case-82": "callback", "case-91": "refunds"}
assert overlap_matrix(previous, current)[("retry", "callback")] == 2
Performance and operating cost
The overlap pass is O(n) time and O(p) space for n anchor tickets and p observed old-new topic pairs. A dense matching algorithm over all old and new topics can cost more, but the reviewer burden is often larger. Keep the anchor set bounded and representative; a tiny set gives fast but unstable alignment. Report its support count and unresolved rate with every trend chart.
Common Mistakes
- Plotting topic ID 4 across retrains as one continuous issue.
- Forcing one-to-one alignment when a theme has split.
- Calling a changed assignment distribution a customer trend without checking traffic mix.
- Hiding unresolved alignments from the dashboard denominator.
Read next
- Topic discovery: corpus boundaries, model choice and human labels
- Project: publish a reviewed support-theme monitor
- Text validation: split conversations, duplicates and time together
- Script profiles and code-switching boundaries in text intake
- Text embeddings: pair labels, hard negatives and versioned vectors
