Self-training proposes labels for an allowed unlabeled training pool, then adds only predictions that meet a declared confidence rule; final evaluation rows remain untouched.
Pseudo-label selection and contamination control
Keep an unlabeled pool inside development
A dispatch model has many shipment records whose handoff outcomes are not yet reviewed. A classifier fitted on genuine training labels scores an unlabeled development pool. Selected high-confidence scores may become pseudo-labels for another fit, but they are still model guesses. The final test cohort cannot enter that pool even without labels: using its features to tune the procedure changes the evaluation question. The split boundary applies to unlabeled records too.
Require trustworthy scores before a confidence rule
A threshold such as 0.9 has no useful meaning if the score is poorly calibrated. Validate probability behavior on disjoint, genuinely labeled development cases before selecting pseudo-labels. The code accepts a frozen score list and returns proposals; it does not train a classifier or claim that the labels are true. It caps proposals per class so one easy majority class cannot consume the entire budget. Calibration is a prerequisite, not a guarantee.
Track provenance through every iteration
Store record ID, original score, model version, threshold and pseudo-label generation round. Do not overwrite an eventual human or mature outcome with the earlier guess. When a guessed label is wrong, repeated self-training can reinforce that error. Audit a sample of proposals by class and site before expanding the procedure. Rare-event prevalence makes majority-class confidence especially deceptive.
Compare against the supervised baseline
A larger training table is not automatically a better model. Fit a supervised-only baseline and a candidate with pseudo-labels under the same development protocol, then compare on genuine later labels. Report which sites contributed pseudo-labels and whether improvement remains after removing weak or correlated cases. If the pseudo-labeled candidate merely becomes more confident without better log loss or policy value, reject it. The triage review supplies the scorecard.
Respect delayed outcomes
Some apparently unlabeled shipments are simply still in transit. A late handoff label may mature after a fixed window, making those cases better held for ordinary supervised learning than guessed now. Keep pseudo-label selection separate from the delayed-label monitor and from human review queues. Outcome maturity determines when a real label can replace a provisional one.
Implementation
# IDs and scores come from an approved unlabeled development pool.
unlabeled_scores = [
("S201", 0.96), ("S202", 0.93), ("S203", 0.58),
("S204", 0.04), ("S205", 0.08), ("S206", 0.91),
]
sealed_test_ids = {"T901", "T902"}
def select_pseudo_labels(scored_rows, low=0.10, high=0.90, per_class=2):
if not 0 <= low < 0.5 < high <= 1:
raise ValueError("invalid confidence limits")
proposals = {0: [], 1: []}
for shipment_id, missed_probability in scored_rows:
if not 0 <= missed_probability <= 1:
raise ValueError("invalid probability")
if missed_probability <= low:
proposals[0].append((shipment_id, 0, missed_probability))
elif missed_probability >= high:
proposals[1].append((shipment_id, 1, missed_probability))
proposals[0].sort(key=lambda row: row[2])
proposals[1].sort(key=lambda row: -row[2])
return proposals[0][:per_class] + proposals[1][:per_class]
selected = select_pseudo_labels(unlabeled_scores)
assert len(selected) == 4
assert not {row[0] for row in selected} & sealed_test_ids
assert {row[1] for row in selected} == {0, 1}Performance and operating cost
Scanning N scored rows costs O(N) time; sorting eligible proposals costs O(N log N) time and O(N) memory. Retraining on a larger set adds model-specific cost. Each iteration also needs human audit, provenance storage and a fresh evaluation against genuine labels.
Common Mistakes
- Do not add sealed-test features to the self-training pool.
- Do not treat a model score as a verified label.
- Do not use a high score cutoff without checking calibration and class support.
Read next
- Weak-label rules, coverage and conflicts
- Active learning with uncertainty and diversity
- Probability calibration: test whether risk scores mean what they say
- Label budget and incremental learning project
Continue the workflow: Nonnegative PU risk objective and its assumptions.
