Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Active-learning pool eligibility and round snapshots

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A query strategy can select only records that are legal to label, unresolved, available at the decision clock and separate from the fixed evaluation set.

Define the label request

An invoice operations model flags documents that require a tax-code review. Each acquired label asks an analyst whether the invoice truly needs specialist review under a versioned policy. A record with no confirmed outcome is not automatically a negative. The label request must include the allowed evidence and a reason code so the analyst can adjudicate it consistently.

Freeze the eligible pool

At each acquisition round, save a pool snapshot with document IDs, customer group, arrival time, policy version and exclusion reason. Remove already labeled invoices, duplicates, legally restricted records and the untouched test set before scoring. A later correction to the pool must create a new version. Otherwise a reported label-efficiency curve cannot be reproduced.

Keep entities together

Several invoices can share a purchase order or supplier template. If a template appears in both the acquisition pool and final test, apparent generalization may be recognition of repeated form. Split by supplier or template family when that matches the deployment claim. Group and time validation establishes the boundary.

Track selection probability

A deterministic top-uncertainty list concentrates labels near a model boundary. That may train a useful classifier, but it is not a representative sample of all invoices. Store the acquisition score, policy and any randomized component. Selection bias affects population claims.

Reserve a random lane

A small random slice can reveal new invoice patterns and provide a training comparison. The final evaluation set must be sampled independently before active selection or analyzed with an appropriate design. The example filters eligible records; it does not choose labels or estimate population performance.

Implementation

python
invoice_pool = [
    {"invoice": "bill-47", "labeled": False, "restricted": False,
     "test_holdout": False, "duplicate_of": None},
    {"invoice": "bill-62", "labeled": True, "restricted": False,
     "test_holdout": False, "duplicate_of": None},
    {"invoice": "bill-83", "labeled": False, "restricted": False,
     "test_holdout": True, "duplicate_of": None},
    {"invoice": "bill-94", "labeled": False, "restricted": False,
     "test_holdout": False, "duplicate_of": "bill-47"},
]

def eligible_invoices(records):
    return [record["invoice"] for record in records
            if not record["labeled"] and not record["restricted"]
            and not record["test_holdout"] and record["duplicate_of"] is None]

assert eligible_invoices(invoice_pool) == ["bill-47"]

Performance and operating cost

Filtering N records costs O(N) time and O(N) output storage. Exact duplicate checks are cheap with stable IDs; near-duplicate templates need indexed similarity review. Persisting each round’s snapshot adds storage but is necessary to reconstruct which invoices could have been queried.

Common Mistakes

  • Do not let the final test set enter the acquisition pool.
  • Do not interpret selected labels as representative prevalence.
  • Do not silently rescore a changed pool under the old round ID.

Read next

ai-data
machine-learning
Storage details