Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Active-learning evaluation and stopping by value

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A label-acquisition strategy earns its place only when it improves a fixed deployment metric at comparable cost on an evaluation set independent of the selected pool.

Freeze evaluation before acquisition

Use a later-period, supplier-grouped test set representative of the invoices the service will face. The active learner must not query it, choose thresholds from it or see its labels between rounds. A selected-label validation set overrepresents uncertain invoices and cannot establish population performance without an explicit sampling design.

Compare cost curves

At each cumulative staff-hour checkpoint, train the active and random-sampling baselines under the same architecture, label policy and development protocol. Plot false tax-code clearances, review volume and calibration with support counts. Compare paired predictions on the same test invoices; a curve drawn against label count hides expensive cases. Learning curves frame the comparison.

Inspect acquisition drift

Query scores change as the model changes. A strategy can move from genuinely informative examples to near duplicates, annotation disputes or an overrepresented supplier. Log pool coverage, novel templates, class mix and adjudication share each round. Oracle disagreement can become the limiting cost.

Stop for a reason

A practical stop rule can combine no material gain over random selection across several budget checkpoints, exhausted staff minutes, rising unresolved-label fraction, or a breached review-capacity limit. Predeclare the rule and use development measurements for operational stopping. Do not repeatedly peek at final-test scores until one looks favorable.

Keep a release path

If active selection helps a narrow slice but harms another, retain the random-lane labels and consider a revised batch policy. After release, track the same decision costs under drift and retraining. The project makes the gate explicit.

Implementation

python
budget_checkpoints = [
    {"hours": 18, "active_missed": 15, "random_missed": 17},
    {"hours": 34, "active_missed": 13, "random_missed": 14},
    {"hours": 51, "active_missed": 13, "random_missed": 12},
]

def no_recent_gain(checkpoints, window):
    if len(checkpoints) < window:
        return False
    recent = checkpoints[-window:]
    return all(row["active_missed"] >= row["random_missed"]
               for row in recent)

assert no_recent_gain(budget_checkpoints, 1) is True
assert no_recent_gain(budget_checkpoints, 2) is False

Performance and operating cost

Checking R budget checkpoints costs O(R) at most and O(1) extra memory. The real expense is parallel random-baseline labeling, repeated model training and a fixed independently labeled test. Those costs are part of proving label efficiency; omitting them can make an acquisition policy look cheaper than it is.

Common Mistakes

  • Do not evaluate population quality on the selected training pool.
  • Do not compare by acquired-label count when label costs vary.
  • Do not repeatedly tune against the untouched final test.

Read next

ai-data
machine-learning
Storage details