An independent adjudicated sample can measure true outcome errors; if it is stratified, its selection weights and denominators must be retained.
Evaluate PU models with an adjudicated sample
Sample beyond easy positives
Review both confirmed-ticket and unresolved pumps from a later period under one failure rubric. Sample across model-score bands, depots, equipment families and observation maturity, including low-score cases. If only suspicious high-score pumps are adjudicated, the measured precision does not represent the deployment population. The target contract defines who was eligible.
Log inclusion probabilities
A stratified audit may oversample high-risk pumps to find enough failures. Keep each sampled record’s chance of selection from its stratum. Weighted totals can recover population rates when every relevant stratum has positive inclusion probability. The code computes a toy weighted recall and states the verified-positive denominator explicitly; it is not an uncertainty interval.
Choose the right metrics
At a fixed review capacity, measure verified failure recall, false review volume and precision against adjudicated truth. Break them out by depot and equipment age. Do not use ticket labels as test truth. If some outcomes remain unresolvable, report that fraction and bound or defer the metric instead of treating unresolved as negative. Rare-event metrics explain the stakes.
Protect the test from model selection
Estimate capture scenarios, fit the model and choose the action threshold on training/development periods. Evaluate one selected candidate on the later adjudicated test. If results lead to revision, create a fresh later test rather than recycling the same cases indefinitely. The final-test boundary still applies.
Report sampling variance and support
Large inverse inclusion weights and few verified failures can make a point estimate unstable. Report sampled and estimated population denominators, effective sample size and uncertainty by site where feasible. If a depot has zero audited failures, do not print a perfect recall. Release review can hold until support improves.
Implementation
audited = [
{"asset": "seal-47", "true_failure": 1, "reviewed": 1, "inclusion": 0.5},
{"asset": "seal-62", "true_failure": 1, "reviewed": 0, "inclusion": 0.25},
{"asset": "seal-83", "true_failure": 0, "reviewed": 1, "inclusion": 0.5},
]
def weighted_failure_recall(records):
if any(not 0 < record["inclusion"] <= 1 for record in records):
raise ValueError("positive inclusion probabilities required")
found = sum(record["true_failure"] * record["reviewed"] / record["inclusion"]
for record in records)
positives = sum(record["true_failure"] / record["inclusion"] for record in records)
if positives == 0:
return None
return found / positives
assert round(weighted_failure_recall(audited), 3) == 0.333
assert weighted_failure_recall([audited[2]]) is NonePerformance and operating cost
Weighted scoring is O(N) time and O(1) auxiliary memory for N audited cases. Adjudication is the dominant cost. High weights inflate uncertainty; a weighted point estimate from three records, as here, is only an arithmetic illustration and must not drive release.
Common Mistakes
- Do not test against confirmation tickets as if they were complete truth.
- Do not forget sampling weights after oversampling suspicious cases.
- Do not call zero observed failures perfect recall.
