Calibrate a versioned prediction-set cutoff, measure minority coverage and stop a release that overwhelms reviewers.
Project: release parcel prediction sets with a bounded review queue
Freeze the parcel frame
The warehouse scanner has four damage classes. Split images by shipment, fit the base model on training shipments and reserve separate calibration and final evaluation shipments. Pin camera firmware, image preprocessing, class order, model digest and scoring formula. Calculate a finite-sample cutoff for the target error rate using only calibration scores. The contract rejects any serving pair with a different model or cutoff revision.
Find the easy-to-miss failure
The candidate passes overall coverage on the final set, but wet cartons from a recently installed camera are rarely included. Full-class sets increase after its lens update. Do not describe marginal coverage as a guarantee for the new camera. Hold automated clearance for that cohort and collect an independent audit sample; reviewed-only labels would bias the estimate. Slice gates require a separate decision for the minority failure.
Test review load and response behavior
Map singleton clear sets to automatic clearance only under the approved policy. Route ambiguous and full sets to manual inspection. A wider candidate cutoff raises daily reviews from 392 to 514 while staffing can handle 470. Hold promotion even if coverage improves. Verify the API returns class names and a set revision, rather than a class-index list that an old client might misread. Coverage and load monitoring should observe both effects.
Deliver a reversible release packet
Record calibration IDs, finite-sample rank, test coverage, per-camera counts, set-width distribution, reviewer forecast and rollback pointer. Shadow the candidate first. If a later camera cohort shifts, route it to review and fit a new cutoff on a fresh approved calibration frame; do not quietly alter the live cutoff. Link decisions to promotion evidence and mature outcomes.
Implementation
def parcel_set_release(report, limits):
if not report["split_disjoint"] or not report["tuple_matched"]:
return "hold:contract"
if report["minority_coverage"] < limits["minority_coverage"]:
return "hold:minority-coverage"
if report["daily_reviews"] > limits["daily_reviews"]:
return "hold:review-capacity"
return "canary"
limits = {"minority_coverage": 0.82, "daily_reviews": 470}
report = {"split_disjoint": True, "tuple_matched": True,
"minority_coverage": 0.88, "daily_reviews": 514}
assert parcel_set_release(report, limits) == "hold:review-capacity"
assert parcel_set_release({**report, "daily_reviews": 392}, limits) == "canary"
Performance and operating cost
The aggregate release gate is O(1) time and space. Calibration sorting is O(n log n) for n examples, while per-request set construction is O(c) for c classes beyond base inference. The expensive part may be reviewer time: extra full sets convert inference uncertainty into a daily human queue.
Common Mistakes
- Using shipment frames from one parcel across training and calibration.
- Promoting from marginal coverage while a new camera slice fails.
- Ignoring the review-volume increase caused by wider sets.
- Sending bare class indices to clients with a different class order.
Read next
- Prediction-set releases: version calibration scores and assumptions
- Prediction-set monitoring: coverage lag, set width and review load
- Slice quality gates when labels are sparse or delayed
- Promotion evidence: bind evaluation, contract and rollback to one digest
- Prediction-outcome joins: evaluate only mature, matched decisions
