Positive–unlabeled learning distinguishes a true event from whether that event was confirmed; an unresolved record is not a verified negative.
Positive–unlabeled target and observation state
Separate truth from observation
A repair network records confirmed seal failures after an inspection. Many pumps have no failure ticket because the follow-up is incomplete or no technician has adjudicated the case. Let Y mean the true 37-day failure event and S mean a confirmed positive ticket. S equals one only when Y equals one; S equals zero can contain either truth value. A classifier trained with S as Y answers the ticket-confirmation question, not automatically the failure question.
Fix the population and clock
Include pumps eligible for scoring at the same inspection decision. Store inspection time, observation cutoff and follow-up maturity. A pump with only 12 observed days cannot be declared free of a 37-day event. If the missing status comes from incomplete follow-up, a time-to-event formulation may be more appropriate than a generic PU classifier. Right censoring describes that case.
Keep label provenance
Record how a positive was confirmed: workshop teardown, sensor alarm review, warranty claim or audit. These processes may select different kinds of failures. The selection method is an input to model validity, not just an annotation note. Selection audit tests whether confirmed positives resemble all positives.
Reserve an adjudicated test
Hold out later inspections and independently resolve a sampled set of their outcomes under a common rubric. Neither PU training nor the rule for selecting a threshold should use this test set. Report unresolved cases separately instead of silently assigning them zero. Evaluation defines its denominators.
State the operating action
The product may send the highest-risk pumps to manual review under a fixed capacity. The goal is to find failures without overwhelming technicians, not to maximize the number of predicted tickets. Review capacity and the project connect scores to actions.
Implementation
inspections = [
{"asset": "seal-47", "confirmed_failure": True, "matured_days": 42},
{"asset": "seal-62", "confirmed_failure": False, "matured_days": 41},
{"asset": "seal-83", "confirmed_failure": False, "matured_days": 12},
]
def observation_state(record, horizon_days):
if record["confirmed_failure"]:
return "labeled-positive"
if record["matured_days"] < horizon_days:
return "unresolved-follow-up"
return "unlabeled-mature" # Still not an adjudicated negative.
states = [observation_state(record, 37) for record in inspections]
assert states == ["labeled-positive", "unlabeled-mature", "unresolved-follow-up"]
assert states.count("labeled-positive") == 1Performance and operating cost
State assignment costs O(N) time for N inspections and O(N) memory when retained. Mature follow-up does not prove absence if the capture system misses failures. Adjudication, provenance logging and outcome lag dominate the operational cost; the code checks the contract only.
Common Mistakes
- Do not relabel every unconfirmed case as a true negative.
- Do not confuse ticket prediction with failure prediction.
- Do not ignore immature follow-up when defining the training population.
