Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Diagnostic performance: separate sensitivity, specificity and predictive value

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A flagger’s sensitivity and specificity describe conditional error rates; the meaning of a positive flag also depends on prevalence.

Define truth and the target frame

A claims audit rule flags duplicate reimbursements for human review. True status comes from a completed independent audit, not from whether the rule raised a flag. Define eligible claims, adjudication window and unit of analysis. If only flagged claims receive verification, false negatives and prevalence are not observed for the full frame. Audit sampling must include some unflagged cases before deployment-wide performance can be estimated.

Keep conditional rates distinct

Sensitivity is the share of true duplicates flagged; specificity is the share of true nonduplicates left unflagged. Positive predictive value is the share of flagged claims that truly duplicate. Holding sensitivity and specificity fixed, positive predictive value falls when duplicates are rarer. The small calculation below shows that a flagger with seemingly strong conditional rates can still send mostly false alerts to reviewers at low prevalence. Rare-event uncertainty matters for every count behind those percentages.

Report the confusion counts

Publish true positives, false positives, false negatives and true negatives with denominators, sampling design and interval method. When audit sampling oversamples flagged claims, raw sample precision is not the live positive predictive value; reweight to the target frame under an approved design. Check groups and time windows, because changes in duplicate prevalence can alter reviewer yield without changing the conditional error rates. Stratified estimation connects the sample to the operational population.

Set an action threshold

A review team cares about missed duplicates and the capacity spent on false alarms. Predeclare acceptable miss and alert rates, then evaluate on mature independent audits. Do not claim a diagnostic rule is accurate from one headline accuracy number when most claims are legitimate. Reviewer agreement checks the quality of the adjudication labels; the project separates label disagreement from model error.

Implementation

python
def positive_flag_value(sensitivity, specificity, prevalence):
    if not all(0 <= value <= 1 for value in
               (sensitivity, specificity, prevalence)):
        raise ValueError("rates must lie between zero and one")
    true_positive_share = sensitivity * prevalence
    false_positive_share = (1 - specificity) * (1 - prevalence)
    flagged_share = true_positive_share + false_positive_share
    return None if flagged_share == 0 else true_positive_share / flagged_share

value = positive_flag_value(0.90, 0.95, 0.02)
assert round(value, 6) == round(0.018 / 0.067, 6)
assert value < 0.30

Performance and operating cost

The probability conversion is O(1) time and space. Reliable estimation requires audited true labels across flagged and unflagged claims and can be expensive. A high reported sensitivity from selectively reviewed claims may be cheaper to obtain but cannot support deployment-wide performance claims.

Common Mistakes

  • Calling sensitivity the probability that a positive flag is correct.
  • Ignoring prevalence when forecasting reviewer workload.
  • Treating unreviewed claims as known true negatives.
  • Reporting only accuracy for a rare duplicate outcome.

Read next

Continue the workflow: Composition reversal: reconcile crude and within-stratum comparisons.

Continue the workflow: Binary outcomes: report absolute risk, risk ratio, and odds on their own scales.

Continue the workflow: Outcome misclassification: correct an observed rate only under stated label-error assumptions.

ai-data
applied-statistics
Storage details