Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Evaluate PU models with an adjudicated sample

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An independent adjudicated sample can measure true outcome errors; if it is stratified, its selection weights and denominators must be retained.

Sample beyond easy positives

Review both confirmed-ticket and unresolved pumps from a later period under one failure rubric. Sample across model-score bands, depots, equipment families and observation maturity, including low-score cases. If only suspicious high-score pumps are adjudicated, the measured precision does not represent the deployment population. The target contract defines who was eligible.

Log inclusion probabilities

A stratified audit may oversample high-risk pumps to find enough failures. Keep each sampled record’s chance of selection from its stratum. Weighted totals can recover population rates when every relevant stratum has positive inclusion probability. The code computes a toy weighted recall and states the verified-positive denominator explicitly; it is not an uncertainty interval.

Choose the right metrics

At a fixed review capacity, measure verified failure recall, false review volume and precision against adjudicated truth. Break them out by depot and equipment age. Do not use ticket labels as test truth. If some outcomes remain unresolvable, report that fraction and bound or defer the metric instead of treating unresolved as negative. Rare-event metrics explain the stakes.

Protect the test from model selection

Estimate capture scenarios, fit the model and choose the action threshold on training/development periods. Evaluate one selected candidate on the later adjudicated test. If results lead to revision, create a fresh later test rather than recycling the same cases indefinitely. The final-test boundary still applies.

Report sampling variance and support

Large inverse inclusion weights and few verified failures can make a point estimate unstable. Report sampled and estimated population denominators, effective sample size and uncertainty by site where feasible. If a depot has zero audited failures, do not print a perfect recall. Release review can hold until support improves.

Implementation

python
audited = [
    {"asset": "seal-47", "true_failure": 1, "reviewed": 1, "inclusion": 0.5},
    {"asset": "seal-62", "true_failure": 1, "reviewed": 0, "inclusion": 0.25},
    {"asset": "seal-83", "true_failure": 0, "reviewed": 1, "inclusion": 0.5},
]

def weighted_failure_recall(records):
    if any(not 0 < record["inclusion"] <= 1 for record in records):
        raise ValueError("positive inclusion probabilities required")
    found = sum(record["true_failure"] * record["reviewed"] / record["inclusion"]
                for record in records)
    positives = sum(record["true_failure"] / record["inclusion"] for record in records)
    if positives == 0:
        return None
    return found / positives

assert round(weighted_failure_recall(audited), 3) == 0.333
assert weighted_failure_recall([audited[2]]) is None

Performance and operating cost

Weighted scoring is O(N) time and O(1) auxiliary memory for N audited cases. Adjudication is the dominant cost. High weights inflate uncertainty; a weighted point estimate from three records, as here, is only an arithmetic illustration and must not drive release.

Common Mistakes

  • Do not test against confirmation tickets as if they were complete truth.
  • Do not forget sampling weights after oversampling suspicious cases.
  • Do not call zero observed failures perfect recall.

Read next

ai-data
machine-learning
Storage details