Build a mature-label report, spot review-selection bias and test a calibrator with a bounded review queue.
Project: monitor claims probabilities and release a new calibrator
Assemble the scoring ledger
The service scores repair claims and routes high-risk cases to inspectors. Capture model, calibrator and threshold revisions with each claim ID, event time, probability, route and acquisition channel. Join only outcomes whose inspection window has closed. Report how many scored claims were eligible, how many were labeled and how many are still pending. The monitoring ledger keeps incomplete labels out of the denominator.
Expose the biased sample
Most inspected claims have outcomes, while auto-cleared claims rarely do. A naive calibration table uses only inspected claims and claims the middle-score band is underconfident. Add an independently selected audit sample of auto-cleared claims and report the two frames separately. One small channel has too few mature labels for a stable band estimate; mark it insufficient instead of drawing a confident line through it. Review corrected labels in a versioned record.
Test a candidate mapping
Fit a candidate calibrator on a separate earlier period. Hold the latest evaluation frame untouched until the mapping is fixed. Compare probability error and band gaps, then rerun the routing threshold on both candidate and incumbent probabilities. The candidate improves global probability error but would send 514 claims per day to a team that can handle 470. Hold it. The release gate requires both probability quality and operating capacity.
Deliver a decision record
Report score-band counts, maturity, audit coverage, slice gaps, review-volume forecast, fit and evaluation windows, and the exact decision tuple. If another candidate passes, shadow it and canary with a rollback pointer. Recheck mature outcomes after deployment; do not treat early review yield as the final outcome. Link the handoff to outcome maturity and threshold policy.
Implementation
def claims_release(candidate, prior, daily_capacity):
if candidate["fit_window"] == candidate["evaluation_window"]:
return "hold:overlap"
if candidate["daily_review_count"] > daily_capacity:
return "hold:capacity"
if candidate["mature_brier"] >= prior["mature_brier"]:
return "hold:no-quality-gain"
return "canary"
prior = {"mature_brier": 0.18}
candidate = {"fit_window": "weeks-47-54", "evaluation_window": "weeks-55-58",
"daily_review_count": 514, "mature_brier": 0.16}
assert claims_release(candidate, prior, 470) == "hold:capacity"
assert claims_release({**candidate, "daily_review_count": 392},
prior, 470) == "canary"
Performance and operating cost
The release check is O(1). The ledger join and band aggregation scale linearly with scored claims if indexed by claim ID; retrospective audit labeling and parallel shadow scoring add cost. Capacity is measured in reviewed claims per day, not merely CPU time, so a faster scorer can still create an unusable release.
Common Mistakes
- Marking pending claims as clean to fill calibration bins.
- Using only inspected claims to describe every scored claim.
- Choosing a calibrator on the protected evaluation period.
- Promoting a better probability score despite exceeding reviewer capacity.
Read next
- Calibration monitoring: pair score bands with matured outcomes
- Calibrator releases: requalify thresholds after score remapping
- Prediction-outcome joins: evaluate only mature, matched decisions
- Selective labels: measure what the model never lets reviewers see
- Score thresholds are release policy, not model metadata
