Labeling only the least-confident tickets can waste reviewers on duplicates and miss confident systematic errors. Balance uncertainty with coverage.
Active learning for text: uncertainty, diversity and coverage
State the acquisition goal
The goal is not to find hard examples in the abstract. It may be to improve rare intent recall, identify new entity forms or reduce multilingual routing errors. Choose a sampling unit—conversation, document or span—and a review budget. Keep the final audit set out of acquisition. Store model version, score, selected reason and eligible pool snapshot so later evaluation can explain why a case entered training.
Mix sampling strategies
Uncertainty sampling surfaces borderline cases, but a model can be confidently wrong on an unseen product. Add random samples to estimate population error and diversity samples to cover new channels, languages and features. Group near duplicates before selection; otherwise a large outage can fill the queue with almost identical tickets. Within a language slice, reserve capacity for rare but consequential labels. Slice evaluation identifies where coverage is weak.
Control the feedback loop
Reviewers should not see the model label as an unquestioned answer. Capture independent decisions for a small overlap set, adjudicate disagreements and log uncertainty. Do not feed a corrected label into training before its policy version and source consent are recorded. A queue full of difficult ambiguous cases may raise disagreement without improving a deployed model. Measure marginal error reduction against a random-sampling baseline after each batch.
Keep evaluation independent
Split by conversation and time before selecting review candidates. Evaluate on a frozen recent holdout that acquisition cannot alter. Report error by slice and review hours consumed, not merely overall accuracy after adding more labels. Disagreement review decides whether a sample has a usable label, and the project implements the queue.
Implementation
def select_review_cases(candidates, quota):
if quota < 1:
raise ValueError("review quota must be positive")
selected, seen_conversations = [], set()
ranked = sorted(candidates, key=lambda row: (-row["priority"], row["case_id"]))
for candidate in ranked:
if candidate["conversation_id"] in seen_conversations:
continue
selected.append(candidate["case_id"])
seen_conversations.add(candidate["conversation_id"])
if len(selected) == quota:
break
return selected
pool = [{"case_id": "case-47", "conversation_id": "thread-k", "priority": 0.93},
{"case_id": "case-48", "conversation_id": "thread-k", "priority": 0.89},
{"case_id": "case-82", "conversation_id": "thread-m", "priority": 0.81}]
assert select_review_cases(pool, 2) == ["case-47", "case-82"]
Performance and operating cost
Sorting n candidates is O(n log n) time and O(n) output or sort space; one pass then enforces conversation diversity. Scoring all unlabeled cases may cost far more than selection itself, so batch or sample the pool with a recorded policy. The scarce resource is reviewer time. Compare improvement per reviewed hour against random selection, and retain random audit samples to detect confident blind spots.
Common Mistakes
- Filling the queue with duplicates from one conversation.
- Selecting only uncertain cases and missing confident model errors.
- Training on final audit examples after reviewers label them.
- Treating an adjudication dispute as a clean training label.
Read next
- Adjudicate text labels before they become training truth
- Project: build a versioned text-label acquisition queue
- Text classification evaluation: inspect slices and allow abstention
- Text validation: split conversations, duplicates and time together
- Text corpus contracts: identity, label timing and annotation rules
Continue the workflow: Adjudicate text labels before they become training truth.
Continue the workflow: Low-resource augmentation: provenance, label drift and slice audit.
Continue the workflow: Text feedback loops: correction provenance and shadow audits.
