Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Active-learning batch selection under annotation cost

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An acquisition batch should trade uncertain cases against duplicate coverage, slice coverage and reviewer time; selecting the highest score repeatedly can waste the budget.

Choose a measurable objective

For tax-code review, uncertainty may be distance from a calibrated decision threshold. It is a ranking signal, not expected improvement by itself: a poorly calibrated model can be confidently wrong, and an unreadable scan can stay uncertain regardless of how many similar scans are labeled. Uncertainty and diversity provide the starting vocabulary.

Reject redundancy within a batch

If 30 invoices use the same supplier form, a top-score list might spend the entire annotation budget on near copies. Enforce a per-template or supplier cap, then compare coverage by customer sector and document quality. A cap can also exclude a common high-risk group, so inspect the selected distribution after each round rather than treating diversity as an automatic improvement.

Use actual label cost

A one-page typed invoice may take 34 seconds to label; a nine-page scan may take several minutes and two reviewers. Rank candidate value against expected review time when budget is measured in staff hours. The code greedily selects by an illustrative value-per-minute score while respecting a template cap. It is a policy sketch, not a claim of globally optimal acquisition.

Keep random comparison stable

For the same staff-minute budget, label an independently sampled random batch with the same policy and reviewer process. Evaluate both resulting models on the same untouched test. More queried records is not necessarily better when one strategy receives easier labels.

Log why a case was picked

Persist model version, score, expected minutes, template ID, pool version and final decision. A policy update or label-cost underestimate can then be diagnosed. Round comparison relies on those logs.

Implementation

python
candidates = [
    {"invoice": "bill-47", "template": "supplier-a", "value": 0.84, "minutes": 2},
    {"invoice": "bill-58", "template": "supplier-a", "value": 0.91, "minutes": 2},
    {"invoice": "bill-76", "template": "supplier-b", "value": 0.72, "minutes": 3},
    {"invoice": "bill-95", "template": "supplier-c", "value": 0.39, "minutes": 1},
]

def choose_batch(records, minute_budget):
    chosen, used_templates, spent = [], set(), 0
    ranked = sorted(records, key=lambda row: row["value"] / row["minutes"], reverse=True)
    for record in ranked:
        if record["template"] in used_templates or spent + record["minutes"] > minute_budget:
            continue
        chosen.append(record["invoice"])
        used_templates.add(record["template"])
        spent += record["minutes"]
    return chosen, spent

assert choose_batch(candidates, 5) == (["bill-58", "bill-95"], 3)

Performance and operating cost

Sorting N candidates costs O(N log N) time and O(N) memory, followed by O(N) selection. Greedy selection may leave budget unused and can miss a better combination. Full model scoring, duplicate detection and human labeling often dominate; compare end-to-end gain per staff hour.

Common Mistakes

  • Do not maximize uncertainty while ignoring repeated templates.
  • Do not compare strategies by label count when label times differ.
  • Do not call the greedy heuristic globally optimal.

Read next

ai-data
machine-learning
Storage details