A probability remapping is a new decision component; it can change who crosses an unchanged numeric threshold.
Calibrator releases: requalify thresholds after score remapping
Treat the calibrator as a release artifact
A claims classifier is unchanged, but a newly fitted calibrator maps raw score 0.47 to probability 0.63 instead of 0.51. If manual inspection starts at 0.60, that claim changes route. Version the calibrator, training labels, fit window, base model and score transformation, then bind them to the threshold policy. Do not roll out a new calibrator as a dashboard-only change. Artifact lineage must identify the exact mapping that produced every decision.
Fit without contaminated evidence
Fit calibration on predictions from observations not used to fit the base model, or use a cross-validated out-of-fold design. Reserve a later untouched evaluation set for the final check. Repeated attempts to choose a mapping on the final set make its estimate optimistic. Time and entity splits still matter: two claims from one repair center can share process artifacts. Snapshot and split replay makes the fit reproducible.
Requalify both probability and workload
Compare reliability bands, a proper probability score, decision precision and recall, and expected manual-review volume on the same eligible frame. Inspect channel and repair-type slices; a better global calibration curve can coexist with a harmful routing change for one segment. Put a capacity limit on new reviews before expanding traffic. Slice gates and threshold change control should approve the combined base-model, calibrator and policy tuple.
Ship a reversible tuple
Shadow both tuples, then canary the new one while retaining the old mapping and threshold. Log the tuple revision with each score and route. If review volume breaches the agreed limit or mature outcomes deteriorate, restore the prior tuple rather than patching the threshold blindly. The reversible pointer allows a single rollback, and the project catches a mapping that looks better overall but overwhelms reviewers.
Implementation
def calibrator_promotion(report, limits):
if report["evaluation_overlap"]:
return "hold:data-overlap"
if report["max_slice_gap"] > limits["max_slice_gap"]:
return "hold:slice-calibration"
if report["review_volume"] > limits["review_capacity"]:
return "hold:review-capacity"
if report["brier_score"] > limits["brier_score"]:
return "hold:probability-quality"
return "canary:bound-tuple"
limits = {"max_slice_gap": 0.08, "review_capacity": 470,
"brier_score": 0.19}
report = {"evaluation_overlap": False, "max_slice_gap": 0.04,
"review_volume": 392, "brier_score": 0.16}
assert calibrator_promotion(report, limits) == "canary:bound-tuple"
assert calibrator_promotion({**report, "review_volume": 514}, limits) == "hold:review-capacity"
Performance and operating cost
The aggregate gate is O(1) time and space. Fitting a calibrator adds a training pass and evaluation cost; shadowing two tuples doubles scoring work for sampled requests. Review capacity is a hard operating constraint because a remapping can move many claims across a threshold without changing model rank order.
Common Mistakes
- Changing a calibrator without versioning the decision tuple.
- Fitting the mapping on base-model training predictions and claiming independent calibration.
- Checking probability error but ignoring manual-review volume.
- Adjusting the threshold after canary without a separate recorded policy revision.
Read next
- Calibration monitoring: pair score bands with matured outcomes
- Project: monitor claims probabilities and release a new calibrator
- Training manifests: link data, code, configuration and artifact
- Score thresholds are release policy, not model metadata
- Promotion evidence: bind evaluation, contract and rollback to one digest
