A prediction set is a release artifact whose coverage claim depends on its calibration frame and scoring rule.
Prediction-set releases: version calibration scores and assumptions
Define what the set means
A warehouse scanner classifies a damaged parcel as crushed, torn, wet or clear. Instead of forcing one class, a prediction-set layer returns every class whose nonconformity score falls within a fitted cutoff. The target is a stated marginal coverage level over data exchangeable with the held-out calibration examples. It is not a guarantee for every camera, damage type or future shifted distribution. Record the base model, class order, nonconformity formula, cutoff, calibration IDs and target error rate together. Slice quality gates still examine the groups that a marginal number can hide.
Keep calibration separate from fitting
Train the base model on its training split. Fit the cutoff on a disjoint calibration split and reserve a later test split for one release check. The cutoff is an empirical quantile of calibration nonconformity scores; use a conservative finite-sample rank rather than interpolating a lower value. Parcel images from one shipment must stay in one split. Otherwise near-duplicate frames can make sets appear smaller and more reliable than they are. Split replay preserves those shipment boundaries.
Pin the whole tuple
Changing preprocessing, class order, calibration source or base model invalidates an old cutoff. Store a digest for each component and reject mixed revisions at serving time. A newer base model with the same labels can change the score distribution. An unchanged cutoff is not automatically reusable. Put the prediction set beside the single-class score and route in the response contract, with an explicit empty-set or full-set fallback. Response contracts keep clients from treating a set as a sorted probability list.
Evaluate usefulness alongside coverage
A set containing all four classes can cover the true class often yet be useless for routing. Measure empirical coverage, singleton rate, average set size, per-camera coverage and the fraction sent to manual review. Include uncertainty for sparse groups. If the widest sets overload reviewers, adjust workflow capacity or model quality; do not silently lower the cutoff after looking at the protected test set. Coverage monitoring pairs the statistical and operational views.
Implementation
from math import ceil
def finite_sample_cutoff(calibration_scores, error_rate):
if not calibration_scores or not 0 < error_rate < 1:
raise ValueError("nonempty scores and an error rate in (0, 1) required")
ordered = sorted(calibration_scores)
rank = ceil((len(ordered) + 1) * (1 - error_rate))
return float("inf") if rank > len(ordered) else ordered[rank - 1]
scores = [0.04, 0.12, 0.18, 0.27, 0.39, 0.48, 0.62, 0.71,
0.76, 0.83, 0.89, 0.93, 0.94, 0.96, 0.97, 0.98,
0.99, 0.995, 0.998, 0.999]
assert finite_sample_cutoff(scores, 0.1) == 0.998
assert finite_sample_cutoff(scores[:4], 0.1) == float("inf")
Performance and operating cost
Sorting n calibration scores costs O(n log n) time and O(n) space. With a small calibration frame, the conservative rank can exceed n and yield an uninformative full set. At serving time, scoring c classes costs O(c) beyond the base model. Review volume grows when sets contain multiple plausible classes.
Common Mistakes
- Claiming per-camera coverage from a marginal target alone.
- Fitting the cutoff on training predictions.
- Reusing a cutoff after changing the model or class order.
- Calling a full-class set a useful answer because it covers the label.
Read next
- Prediction-set monitoring: coverage lag, set width and review load
- Project: release parcel prediction sets with a bounded review queue
- Training replay: freeze the cohort, split and runtime
- Slice quality gates when labels are sparse or delayed
- Inference API contracts: version the decision, not only the payload
