Calibration measures whether predicted probabilities match observed event rates in comparable, mature outcome cohorts.
Calibration monitoring: pair score bands with matured outcomes
Record the probability contract
A claims model predicts the chance that a submitted repair claim needs manual inspection. Store raw score, calibrated probability, calibrator revision, model revision, decision threshold, event time, cohort and eventual label join key. A model can rank claims well while assigning probabilities that are too high or low. Compare observed event rates with mean predicted probabilities within score bands, and publish the band sample size. Threshold policy uses those probabilities but answers a different decision question.
Wait for labels to mature
A claim may be inspected days after its score was issued. Exclude recent predictions whose outcome window has not closed; otherwise a quiet-looking band may simply lack labels. Record both scored count and mature joined count by band and cohort. Late labels and corrected outcomes should update a versioned ledger, not silently replace an old report. Outcome joins define maturity and label corrections keep changes traceable.
Watch selection bias
Claims routed to manual review are more likely to receive confirmed labels than auto-cleared claims. Calibration measured only on reviewed claims does not represent all scored claims. Use an independent audit sample, or report the reviewed population honestly and avoid a site-wide calibration claim. Split by acquisition channel, repair category and score band, then show uncertainty where counts are small. Audit sampling supplies a second observation frame.
Trigger investigation, not automatic retuning
A band whose observed rate drifts away from its mean prediction may reflect prevalence change, data quality, label latency or a broken scoring pipeline. Check calibration error over comparable windows with minimum sample counts and a stable label definition. Investigate cohort composition before refitting. A new calibrator must pass its own holdout and policy review. Calibrator requalification governs the change rather than letting a dashboard rewrite production scores.
Implementation
def band_calibration(records, lower, upper, minimum_mature):
eligible = [record for record in records if record["mature"]
and lower <= record["probability"] < upper]
if len(eligible) < minimum_mature:
return {"state": "insufficient", "count": len(eligible)}
predicted = sum(row["probability"] for row in eligible) / len(eligible)
observed = sum(row["outcome"] for row in eligible) / len(eligible)
return {"state": "measured", "count": len(eligible),
"predicted": predicted, "observed": observed}
claims = [{"probability": 0.42, "outcome": 1, "mature": True},
{"probability": 0.48, "outcome": 0, "mature": True},
{"probability": 0.46, "outcome": 1, "mature": False}]
report = band_calibration(claims, 0.4, 0.5, 2)
assert report["state"] == "measured" and report["count"] == 2
assert report["observed"] == 0.5
Performance and operating cost
One band scan is O(n) time and O(n) temporary space for n records; streaming sums can reduce extra space to O(1). A full monitoring job should compute all bands in one pass. Audit labeling has an operational cost, but a reviewed-only sample may make an inexpensive report misleading.
Common Mistakes
- Treating ranked scores as calibrated probabilities without checking outcomes.
- Counting immature predictions as negative outcomes.
- Reporting calibration from the review queue as if it covered all claims.
- Refitting on the same holdout later used to approve the calibrator.
Read next
- Calibrator releases: requalify thresholds after score remapping
- Project: monitor claims probabilities and release a new calibrator
- Prediction-outcome joins: evaluate only mature, matched decisions
- Selective labels: measure what the model never lets reviewers see
- Score thresholds are release policy, not model metadata
