Skip to content
AITroveRead. Build. Understand.
Make this comfortable

PU class prior and confirmation-rate sensitivity

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Positive prevalence and confirmation probability determine how observed ticket rates translate into true event risk under a stated selection model.

Separate prevalence from capture

The observed positive-ticket fraction is not the true failure prevalence when some failures are never confirmed. Under constant capture c, observed rate equals true prevalence times c. If tickets cover 12 percent of eligible pumps and c is 0.6, the implied prevalence is 20 percent. This arithmetic is conditional on the selection model and matched population; it is not a way to discover c from tickets alone.

Bound the plausible scenarios

Use independently adjudicated audits, process knowledge or a validated capture study to set a range for c. Apply each value to the same observed rate and inspect whether review capacity, threshold choice or site performance changes. The code evaluates three illustrative scenarios; none is presented as an estimated truth. The SCAR audit explains why one c may fail by depot.

Watch impossible combinations

If observed ticket rate exceeds an assumed c, dividing yields a prevalence above one; the assumption is incompatible with that population. Clipping the result to one hides the contradiction. Likewise, an observed score divided by c may exceed one because the fitted score or c is misspecified. Treat that as a model diagnostic, not merely a formatting issue.

Avoid prior leakage

Do not choose prevalence from a later test period and feed it into earlier training without marking the data transfer. If failure prevalence shifts after a maintenance policy change, a historical prior may no longer apply. Prior shift addresses an adjacent deployed-model problem.

Carry uncertainty into the decision

A release decision should show the range of expected review volume and missed failures across plausible c and prevalence values. If one plausible scenario violates the staffing limit, hold the automation or use a more conservative manual review path. The project makes that gate explicit.

Implementation

python
observed_ticket_rate = 0.12
capture_scenarios = (0.45, 0.60, 0.75)

def implied_prevalence(observed_rate, capture_probability):
    if not 0 <= observed_rate <= 1 or not 0 < capture_probability <= 1:
        raise ValueError("valid rates required")
    if observed_rate > capture_probability:
        raise ValueError("observed rate exceeds assumed capture")
    return observed_rate / capture_probability

prevalence_scenarios = [implied_prevalence(observed_ticket_rate, capture)
                        for capture in capture_scenarios]
assert [round(value, 3) for value in prevalence_scenarios] == [0.267, 0.2, 0.16]
assert prevalence_scenarios[0] > prevalence_scenarios[-1]

Performance and operating cost

Evaluating S capture scenarios costs O(S) time and O(S) storage. Obtaining defensible capture bounds is much costlier than the arithmetic because it needs outcome adjudication or a capture study. When the selection model fails, the range is not an uncertainty interval for truth.

Common Mistakes

  • Do not identify true prevalence from ticket rate alone.
  • Do not silently clip an impossible implied prevalence.
  • Do not reuse a future-period prior in an earlier training decision.

Read next

ai-data
machine-learning
Storage details