Create a review queue that balances uncertain, diverse and random cases, captures independent labels and admits only adjudicated records to training.
Project: build a versioned text-label acquisition queue
Define the queue contract
Each candidate carries case identity, conversation group, capture time, channel, reviewed slice if known, model version and selection reason. The final audit set is excluded before scoring. Allocate a budget for uncertainty, coverage of rare slices and random traffic. Group near duplicates by conversation so one event cannot consume the whole batch. Keep raw text visible only to authorized reviewers.
Review and adjudicate
Assign independent reviewers to an overlap sample and collect labels under one explicit policy version. If they disagree, route the case to an adjudicator with the original context and both rationales. Ambiguous or unresolved cases stay out of supervised training until the policy defines their treatment. Disagreement policy governs this step; sampling design governs intake.
Release a dataset revision
Write an immutable manifest of accepted case IDs, source revisions, label policy, reviewers and split assignment. Check that no conversation crosses train and audit and that training excludes withheld or unconsented data. Evaluate the model after each acquisition batch on the same frozen holdout; report improvement per reviewer hour and slices that remain weak. Do not call a larger dataset better if its labels are less consistent.
Watch the workflow
Monitor queue age, disagreement rate, duplicate selection, accepted-label yield, random-sample error and model gains. A highly uncertain pool can be intrinsically ambiguous, so decreasing yield is a signal to revise the taxonomy. Maintain a rollback to the prior dataset and model. Retention and deletion policies must remove raw review text and derived training copies when required.
Implementation
def admit_training_record(review_record):
required = {"case_id", "source_revision", "policy_version", "split", "label"}
if not required <= set(review_record):
raise ValueError("review record lacks lineage")
if review_record.get("state") != "adjudicated" or review_record["split"] != "train":
return False
if not review_record.get("use_permitted", False):
return False
return True
record = {"case_id": "case-47", "source_revision": "v3", "policy_version": "p7",
"split": "train", "label": "refund", "state": "adjudicated", "use_permitted": True}
assert admit_training_record(record)
Performance and operating cost
Admission is O(f) time for f record fields and O(1) extra space. Queue scoring can be costly, but reviewer time and source retention dominate the budget. Track accepted labels per reviewer hour, not just candidates displayed. A strict lineage check adds a small amount of work and prevents silent leakage of audit examples or unpermitted text into training.
Common Mistakes
- Sending held-out audit cases into the acquisition queue.
- Treating a submitted label as adjudicated training truth.
- Omitting policy version or source revision from a dataset export.
- Measuring model gain without accounting for reviewer time and duplicate selection.
Read next
- Active learning for text: uncertainty, diversity and coverage
- Adjudicate text labels before they become training truth
- Text corpus contracts: identity, label timing and annotation rules
- Text validation: split conversations, duplicates and time together
- Project: enforce a privacy-safe support-text pipeline
Continue the workflow: Project: audit ticket families before publishing a text classifier score.
