Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Active-learning queues: uncertainty, coverage and selection bias

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Model uncertainty can prioritize review effort, but selected labels describe the selected queue unless a separate sample supports wider claims.

Choose an acquisition rule

Select candidates by calibrated uncertainty, disagreement among models or expected operational value. Deduplicate near-identical receipts before allocation so a scanner burst does not consume the budget. Reserve slots for random exploration and under-covered sources. Calibration matters if a score is interpreted as uncertainty; low confidence may also reflect a shifted scanner or missing input.

Preserve inclusion information

Log the acquisition score, queue policy, model version, source stratum and probability of selection when randomized. Without those fields, an evaluated model may appear to improve simply because the next training batch contained unusually difficult cases. The final held-out test must remain untouched by the selection loop. The sampling frame limits any population-level performance statement.

Balance learning and auditing

The active queue is built to improve the model. A blinded random audit is built to estimate current quality. Budget and report them separately even if the same reviewers operate both workflows. Review false negatives from production after their labels mature, but do not leak those outcomes into a prospective test whose prediction time preceded them.

Simulate a cycle

A queue has 47 receipts: 20 uncertain, 12 from a new scanner and 15 routine. With 12 review slots, choose six uncertain, three new-scanner and three randomized routine items. Record the rule and inclusion route for every chosen item. Refit only on reviewed labels, then evaluate on a frozen future batch rather than on the acquired queue.

Implementation

python
def select_review_queue(scored_receipts, review_slots):
    ranked = sorted(scored_receipts,
                    key=lambda receipt: (-receipt["uncertainty"], receipt["receipt_id"]))
    chosen = []
    seen_sources = set()
    for receipt in ranked:
        if len(chosen) == review_slots:
            break
        if receipt["source"] not in seen_sources:
            chosen.append(receipt["receipt_id"])
            seen_sources.add(receipt["source"])
    chosen_ids = set(chosen)
    for receipt in ranked:
        if len(chosen) == review_slots:
            break
        if receipt["receipt_id"] not in chosen_ids:
            chosen.append(receipt["receipt_id"])
            chosen_ids.add(receipt["receipt_id"])
    return chosen

Performance and operating cost

Sorting N candidates costs O(N log N) time and O(N) space. The snippet demonstrates a deterministic ranking with source coverage; a real audit needs a separate randomized sampler and logged inclusion probabilities.

Common Mistakes

  • Do not evaluate population quality on an uncertainty-selected queue.
  • Do not let duplicate scans exhaust the review budget.
  • Do not use the final test set to drive acquisition.

Read next

Continue the workflow: Active-learning batch selection under annotation cost.

ai-data
data-annotation
Storage details