Active learning chooses which unlabeled cases receive costly human review, often combining model uncertainty with a coverage rule so the batch is not a stack of near duplicates.
Active learning with uncertainty and diversity
Define the review budget and candidate pool
A dispatch analyst can review three ambiguous shipment records this week. Only development-period unlabeled cases are eligible; no final-test row belongs in the queue. The classifier supplies scores, but they are ranking signals, not verified event rates unless calibrated. For a binary task, scores near one half are one uncertainty proxy. Pseudo-labeling instead assigns provisional labels without asking the analyst.
Prevent one site from taking every slot
Uncertainty alone may pick many nearly identical shipments from the same warehouse. The code sorts by closeness to one half and allows one candidate per site. This simple site quota is a diversity policy, not a full optimal batch selector. A different task may require route, equipment type or shift coverage. Record the chosen rule before inspection so a desired result cannot guide the queue retrospectively.
Expect sampling bias in the reviewed labels
The reviewed set overrepresents uncertain cases by design. Its raw event frequency does not estimate the production prevalence, and accuracy on that set is not an unbiased deployment metric. Keep a separate random or policy-relevant audit sample, plus a sealed future cohort, for evaluation. The prevalence guide explains how class mix affects precision.
Audit reviewer consistency
Some ambiguous cases are ambiguous because the event clock or evidence is unclear. Give reviewers an explicit label definition and an abstain or escalation path. Re-review a small overlap sample to measure disagreement; do not force a label merely because the model requested one. Store reviewer version, decision time and adjudication history. The annotation curriculum provides the wider workflow.
Measure value per label
After adding reviewed cases, retrain the development model and compare it with a random-sampling acquisition baseline under equal label cost. Gains can flatten once a particular boundary is well sampled. Report performance on genuine held-out labels and by site. Feature selection must still happen inside the training folds when the labeled set changes.
Implementation
unlabeled_pool = [
("S401", "Harbor", 0.49), ("S402", "Harbor", 0.51),
("S403", "Inland", 0.47), ("S404", "North", 0.54),
("S405", "South", 0.88), ("S406", "Inland", 0.62),
]
sealed_test_ids = {"T903", "T904"}
def review_queue(scored_cases, budget):
ordered = sorted(scored_cases, key=lambda case: (abs(case[2] - 0.5), case[0]))
selected = []
covered_sites = set()
for shipment_id, site, probability in ordered:
if site in covered_sites:
continue
selected.append((shipment_id, site, probability))
covered_sites.add(site)
if len(selected) == budget:
break
return selected
queue = review_queue(unlabeled_pool, budget=3)
assert len(queue) == 3
assert len({site for _, site, _ in queue}) == 3
assert not {shipment_id for shipment_id, _, _ in queue} & sealed_test_ids
assert queue[0][0] == "S401"Performance and operating cost
Sorting N candidates costs O(N log N) time and O(N) memory, while scoring the pool adds model-specific cost. Human review is the limiting resource. Site quotas may improve coverage but can displace a highly informative case, so compare acquisition policies at equal review budget.
Common Mistakes
- Do not evaluate deployment accuracy on only the actively selected review queue.
- Do not fill the queue with duplicate records from one site.
- Do not force a disputed human label into training without adjudication.
Read next
- Pseudo-label selection and contamination control
- Weak-label rules, coverage and conflicts
- Handoff triage model review project
- Label budget and incremental learning project
Continue the workflow: Hard-negative mining without false-negative shortcuts.
Continue the workflow: Active-learning pool eligibility and round snapshots.
