A classifier can rank classes correctly while expressing unsafe confidence; a held-out temperature and an explicit abstention rule address different parts of that problem.
Temperature scaling and selective risk
Fit confidence after the model
Freeze the trained classifier and collect logits and labels from a separate calibration split. Divide every class logit by one positive temperature and choose the temperature that reduces calibration negative log likelihood. This one-parameter map changes probability sharpness while preserving each example’s class ordering. It cannot fix a wrong top class or a missing input feature. Logits must be saved before softmax so the map is applied in the intended place.
Keep splits independent
The calibration split cannot be the same population used to fit classifier weights if the goal is out-of-sample confidence. Avoid tuning the temperature repeatedly on the final test set, and group images from one physical receipt on one side of each split. A new acquisition device may need a fresh calibration study, but per-device fits require enough observations and a clear fallback. Domain-shift auditing should report probability quality by source, not just average accuracy.
Measure what a probability means
Compare mean confidence with observed correctness in fixed or adaptive bins, and report the sample count per bin. A small expected calibration error can hide rare high-confidence failures because binning and class balance matter. Use negative log likelihood and a proper score alongside diagrams, then inspect high-impact slices. The temperature is positive; parameterizing its logarithm keeps optimization in the valid range. Fit on frozen logits, not on a model that continues to change during calibration.
Add an abstention policy
After calibration, accept a model answer only when its maximum probability clears a selected threshold; otherwise send it for manual review or an alternate workflow. Coverage is the share accepted. Selective risk is the error fraction among accepted cases. Raising the threshold usually lowers coverage but does not guarantee lower risk in every subgroup. Choose the threshold on validation under a stated review-capacity or error-budget rule, then report both coverage and risk on untouched data.
Test the operating limit
Report class-specific coverage, wrong high-confidence predictions and the number of manual reviews per hundred receipts. On an all-blurry batch, a high-confidence model can still be wrong; calibration does not certify that uncertainty reflects all unfamiliar inputs. Keep the model revision, temperature, threshold and label map together. The serving system should reject nonfinite logits and missing preprocessing fields before confidence calculation. Inference contracts protect the output decision.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
calibration_logits = torch.tensor([[2.7, 0.4, -0.8], [0.3, 2.2, -0.2],
[1.8, 1.2, -0.5], [-0.2, 0.6, 1.9],
[1.4, 0.7, -0.3], [0.5, 1.3, 0.2]])
calibration_labels = torch.tensor([0, 1, 1, 2, 0, 1])
log_temperature = nn.Parameter(torch.tensor(0.0))
temperature_optimizer = torch.optim.Adam([log_temperature], lr=0.047)
for _ in range(83):
temperature_optimizer.zero_grad(set_to_none=True)
calibrated_loss = functional.cross_entropy(
calibration_logits / log_temperature.exp(), calibration_labels)
calibrated_loss.backward()
temperature_optimizer.step()
temperature = log_temperature.detach().exp()
calibrated_probabilities = functional.softmax(
calibration_logits / temperature, dim=1)
assert temperature > 0
assert torch.equal(calibration_logits.argmax(dim=1),
calibrated_probabilities.argmax(dim=1))
assert torch.isfinite(calibrated_probabilities).all()Performance and operating cost
Fitting one temperature costs repeated softmax and loss evaluations over the calibration logits, roughly O(iterations × cases × classes), and requires storing logits and labels. Serving adds one scalar division and softmax per prediction. Abstention can reduce automated errors at the price of review workload; that workload is a system cost and must be measured at the selected coverage. This tiny six-case fit illustrates mechanics, not a reliable production calibration.
Common Mistakes
- Do not fit confidence on the untouched final test set.
- Do not expect temperature scaling to change the predicted top class.
- Do not quote selective risk without coverage and subgroup counts.
Read next
- Logits, cross-entropy and gradients: align the training calculation
- Training and validation modes: measure the model you will serve
- Staged unfreezing and domain-shift audits
- Inference contracts: preserve preprocessing and measure tail latency
- Project: audit receipt confidence and review handoff
Continue the workflow: Post-training quantization and calibration contracts.
Continue the workflow: Project: stress-test receipt decisions under bounded image changes.
Continue the workflow: Project: attribute service incidents with a time-safe graph.
Continue the workflow: Quantile forecast loss and crossing policy.
Continue the workflow: Class-aware suppression and detection score thresholds.
Continue the workflow: Project: review telemetry anomalies with an invertible flow.
Continue the workflow: Project: classify machine alarms from short audio windows.
Continue the workflow: Deep ensembles, seed diversity and predictive disagreement.
