A PU release decision must establish that unconfirmed cases remain unlabeled, selection assumptions are defensible, and a separately adjudicated test supports the action threshold.
Project: release review for unresolved pump failures
Freeze the case definition
A service team wants to rank pumps for manual seal inspection. Record the 37-day event definition, eligible inspection cohort, confirmation channels, observation maturity and data cutoffs. S equals a verified ticket; Y is the failure truth. Keep the two states in separate fields. The observation contract is the first release condition.
Audit the labeling process
Compare confirmed and independently adjudicated positives across depots and severity. If capture varies strongly by site, do not apply one constant correction without qualification. Specify plausible confirmation and prevalence ranges, then rerun the capacity decision across them. Selection and prior sensitivity define this evidence.
Compare training choices
Evaluate a ticket-versus-unlabeled baseline, a nonnegative PU objective and a supervised model trained only on fully adjudicated cases. Keep architecture and feature access comparable. The PU objective’s arithmetic depends on the chosen prior; a lower training risk is not the release metric. The risk lesson explains the assumption.
Score the held-out action
Independently adjudicate a later stratified test sample, retain inclusion probabilities, and estimate verified failure recall and false-review burden at the proposed review budget. Report mature-positive counts and site slices. A threshold that helps the pooled estimate but misses one depot’s failures is a hold. Weighted evaluation supplies the denominator.
Require a fallback
If audit support is too thin, keep the previous rule or send uncertain cases to manual review while more outcomes mature. Store the old model, threshold and label-contract version. The code is a gate over an illustrative packet, not a claim that the sample proves field benefit.
Implementation
release_packet = {
"unlabeled_kept_distinct": True,
"selection_audit_complete": True,
"prior_sensitivity_stable": False,
"adjudicated_test_positive_count": 43,
"minimum_test_positives": 38,
"worst_depot_failure_recall": 0.57,
"minimum_depot_recall": 0.64,
"review_capacity_met": True,
"rollback_ready": True,
}
def pu_release_gate(packet):
holds = []
if not packet["unlabeled_kept_distinct"] or not packet["selection_audit_complete"]:
holds.append("label contract")
if not packet["prior_sensitivity_stable"]:
holds.append("prior sensitivity")
if packet["adjudicated_test_positive_count"] < packet["minimum_test_positives"] or packet["worst_depot_failure_recall"] < packet["minimum_depot_recall"]:
holds.append("held-out failure recall")
if not packet["review_capacity_met"] or not packet["rollback_ready"]:
holds.append("operating controls")
return "hold: " + ", ".join(holds) if holds else "eligible for shadow review"
assert pu_release_gate(release_packet) == "hold: prior sensitivity, held-out failure recall"Performance and operating cost
The gate is O(1). Its evidence requires confirmed-ticket review, an adjudicated later-period sample, capture-scenario analysis, repeated model training and a capacity simulation. Those costs should be planned before promising automated review. A passing gate permits a controlled shadow phase, not a claim of complete outcome ascertainment.
Common Mistakes
- Do not promote on confirmation-ticket accuracy.
- Do not suppress a site regression behind a pooled metric.
- Do not call an untested capture assumption a measured fact.
