Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: release review for unresolved pump failures

Last updated: 5 Oct 20265 min read
project
AdvancedBy AITrove Editorial

A PU release decision must establish that unconfirmed cases remain unlabeled, selection assumptions are defensible, and a separately adjudicated test supports the action threshold.

Freeze the case definition

A service team wants to rank pumps for manual seal inspection. Record the 37-day event definition, eligible inspection cohort, confirmation channels, observation maturity and data cutoffs. S equals a verified ticket; Y is the failure truth. Keep the two states in separate fields. The observation contract is the first release condition.

Audit the labeling process

Compare confirmed and independently adjudicated positives across depots and severity. If capture varies strongly by site, do not apply one constant correction without qualification. Specify plausible confirmation and prevalence ranges, then rerun the capacity decision across them. Selection and prior sensitivity define this evidence.

Compare training choices

Evaluate a ticket-versus-unlabeled baseline, a nonnegative PU objective and a supervised model trained only on fully adjudicated cases. Keep architecture and feature access comparable. The PU objective’s arithmetic depends on the chosen prior; a lower training risk is not the release metric. The risk lesson explains the assumption.

Score the held-out action

Independently adjudicate a later stratified test sample, retain inclusion probabilities, and estimate verified failure recall and false-review burden at the proposed review budget. Report mature-positive counts and site slices. A threshold that helps the pooled estimate but misses one depot’s failures is a hold. Weighted evaluation supplies the denominator.

Require a fallback

If audit support is too thin, keep the previous rule or send uncertain cases to manual review while more outcomes mature. Store the old model, threshold and label-contract version. The code is a gate over an illustrative packet, not a claim that the sample proves field benefit.

Implementation

python
release_packet = {
    "unlabeled_kept_distinct": True,
    "selection_audit_complete": True,
    "prior_sensitivity_stable": False,
    "adjudicated_test_positive_count": 43,
    "minimum_test_positives": 38,
    "worst_depot_failure_recall": 0.57,
    "minimum_depot_recall": 0.64,
    "review_capacity_met": True,
    "rollback_ready": True,
}

def pu_release_gate(packet):
    holds = []
    if not packet["unlabeled_kept_distinct"] or not packet["selection_audit_complete"]:
        holds.append("label contract")
    if not packet["prior_sensitivity_stable"]:
        holds.append("prior sensitivity")
    if packet["adjudicated_test_positive_count"] < packet["minimum_test_positives"] or packet["worst_depot_failure_recall"] < packet["minimum_depot_recall"]:
        holds.append("held-out failure recall")
    if not packet["review_capacity_met"] or not packet["rollback_ready"]:
        holds.append("operating controls")
    return "hold: " + ", ".join(holds) if holds else "eligible for shadow review"

assert pu_release_gate(release_packet) == "hold: prior sensitivity, held-out failure recall"

Performance and operating cost

The gate is O(1). Its evidence requires confirmed-ticket review, an adjudicated later-period sample, capture-scenario analysis, repeated model training and a capacity simulation. Those costs should be planned before promising automated review. A passing gate permits a controlled shadow phase, not a claim of complete outcome ascertainment.

Common Mistakes

  • Do not promote on confirmation-ticket accuracy.
  • Do not suppress a site regression behind a pooled metric.
  • Do not call an untested capture assumption a measured fact.

Read next

ai-data
machine-learning
Storage details