A model changes which cases receive labels, so evaluation coverage must include decisions hidden by its own routing policy.
Selective labels: measure what the model never lets reviewers see
Map the observation path
A receipt-risk model sends suspicious cases to human review and releases the rest. Reviewers label many routed cases, while released cases may never receive an immediate fraud outcome. A dashboard built only from reviewed labels describes the selected queue, not all scored receipts. Record each decision, review assignment, label availability and observation maturity as separate events. Outcome maturity handles delay; selective observation asks why a label exists at all.
Reserve an independent audit lane
Define a permitted, small and randomized sample of cases from otherwise unreviewed routes for later adjudication. The sampling probability must be recorded at assignment time, not inferred after outcomes arrive. Keep privacy, reviewer capacity and product policy limits explicit. This lane estimates blind spots without claiming that one sample fixes every bias. Stratify only on pre-decision attributes and retain enough examples near the threshold. Exposure logging keeps the scored population distinct from the audited population.
Report denominator and uncertainty
Show total scored, routed, randomly audited, matured and unresolved counts by cohort. Error rates from voluntary or risk-triggered reviews cannot be generalized to all decisions without assumptions. A random audit estimate may use inverse sampling weights when probabilities vary, but extreme weights inflate variance. Publish the effective sample size and a confidence interval. Slice uncertainty prevents a small audited subgroup from being called a clear pass.
Protect the evaluation lane
Freeze audit assignment before labels are known and keep its evaluation subset separate from training and threshold tuning. A policy change alters both decision quality and which cases become visible; compare policies on a shared independent audit frame or an experiment designed for that purpose. Policy-shift comparison covers this issue, while the project exposes a misleading reviewed-only improvement.
Implementation
def audit_coverage(scored, reviewed, audited, matured):
if scored <= 0 or min(reviewed, audited, matured) < 0:
raise ValueError("invalid counts")
if reviewed > scored or audited > scored or matured > reviewed + audited:
raise ValueError("inconsistent counts")
return {"audit_share": audited / scored,
"observed_share": matured / scored,
"unobserved": scored - matured}
coverage = audit_coverage(4_700, 820, 235, 914)
assert coverage["audit_share"] == 0.05
assert coverage["unobserved"] == 3_786
Performance and operating cost
The count gate is O(1) time and space after aggregation. Random audits consume adjudication time and may carry product or privacy costs; too few audits yield wide intervals, while too many displace routine review. Preserve the assignment probability and maturity window so the measurement can be interpreted later.
Common Mistakes
- Treating reviewed cases as a random sample of all predictions.
- Recording audit membership only after seeing the label.
- Reporting an error rate without the number of unobserved decisions.
- Using the audit holdout to tune the next model and then to approve it.
Read next
- Feedback policy shift: compare models when labels depend on routing
- Project: audit blind spots in receipt-risk feedback
- Prediction-outcome joins: evaluate only mature, matched decisions
- Model experiments: separate assignment from actual exposure
- Slice quality gates when labels are sparse or delayed
Continue the workflow: Project: monitor claims probabilities and release a new calibrator.
Continue the workflow: Cohort release gates: compare harm, coverage and uncertainty together.
