A selective predictor acts automatically only on cases outside a declared uncertainty band and routes the remainder to review, trading automation coverage for error risk and human workload.
Selective prediction and review capacity
Define acceptance and escalation before scoring
A missed-handoff probability below a low cutoff can be treated as routine; above a high cutoff it may trigger an alert; values between the cutoffs go to an analyst. The code uses fixed bounds on already-issued probabilities and reports the fraction handled automatically plus errors on those handled. The bounds are policy settings selected in development, not on the final test labels. Calibration matters if the numeric bounds represent expected risk.
Report coverage with selective error
An automated subset can look accurate because the system escalates every difficult case. Always pair error on accepted cases with the accepted fraction, counts, and outcomes among escalated cases. A policy handling five percent of traffic at zero observed error may still overload reviewers. Report the full population’s realized cost, not only a headline accepted accuracy. Decision costing turns each branch into a resource question.
Measure queue capacity and delay
A human-review branch has arrival rate, service time and a deadline. If the queue cannot process all escalated shipments before dispatch closes, review is not an effective fallback. Set a maximum daily escalation volume and a timeout policy. Consider interval width or a missing-feature flag as an additional escalation trigger when raw probability does not capture uncertainty. Interval coverage may expose the conditions that need review.
Check group effects
A global uncertainty band may send far more cases from one site to manual review, or automate cases for a site where the model is less accurate. Audit acceptance, false negatives and review wait by group with support counts. The group audit is a diagnostic framework, not a guarantee that a chosen band is fair. Document the reason for different treatment if site-specific thresholds are considered.
Keep feedback unbiased
If reviewers label only uncertain cases, their labels do not represent the automatically handled population. Reserve a random audit sample from automated decisions and maintain a later sealed evaluation cohort. Do not retrain from the review queue alone without addressing its selection mechanism. Active learning uses selected labels for development while keeping evaluation separate.
Implementation
# Shipment ID, issued missed-handoff probability, mature outcome.
cases = [("S601", 0.08, 0), ("S602", 0.19, 0),
("S603", 0.43, 1), ("S604", 0.55, 0),
("S605", 0.82, 1), ("S606", 0.94, 1)]
low_cutoff, high_cutoff = 0.25, 0.75
def disposition(probability):
if probability < low_cutoff:
return "routine", 0
if probability > high_cutoff:
return "alert", 1
return "review", None
decisions = [(shipment_id, *disposition(probability), actual)
for shipment_id, probability, actual in cases]
automatic = [row for row in decisions if row[1] != "review"]
reviewed = [row for row in decisions if row[1] == "review"]
automatic_error_rate = sum(predicted != actual
for _, _, predicted, actual in automatic) / len(automatic)
automatic_coverage = len(automatic) / len(decisions)
assert automatic_coverage == 4 / 6
assert automatic_error_rate == 0
assert len(reviewed) == 2Performance and operating cost
Routing N issued scores costs O(N) time and O(1) counters if streamed. Review workload scales with escalated volume; latency and staffing cost can dominate model inference. Store issued decisions and outcomes to audit the entire policy, not only the accepted subset.
Common Mistakes
- Do not report selective accuracy without automation coverage.
- Do not assume human review is available before the decision deadline.
- Do not evaluate future performance only from the selectively reviewed queue.
Read next
- Group error gaps and policy audit
- Prediction interval coverage by operating slice
- Active learning with uncertainty and diversity
- Uncertainty and escalation review project
Continue the workflow: Exploration budget and action guardrails.
