Split conformal regression uses absolute errors from a held-out calibration set to place an empirical margin around a fixed point predictor.
Split conformal intervals for clearance forecasts
Reserve a calibration cohort
A shipment model predicts hours until clearance. Fit the predictor on training shipments, select its settings on development data, then freeze it before scoring a separate calibration cohort. None of those calibration outcomes may have influenced the model or preprocessing. The margin comes from absolute calibration residuals, not training residuals. The split guide explains why related shipments and future timestamps must not leak between roles.
Choose a finite-sample rank
For a requested miscoverage fraction alpha and n calibration residuals, use the rank ceiling of (n plus one) times (one minus alpha), with one-based indexing. When that rank exceeds n, a finite margin cannot support the requested finite-sample claim without a wider convention; the code raises instead of quietly taking the largest observed error. The resulting interval is the point forecast plus or minus the selected residual. It is a teaching implementation of the split procedure.
State the coverage scope
The familiar marginal coverage guarantee relies on calibration and future examples being exchangeable under the prediction procedure. A later warehouse policy, selective labeling or dependent repeated scans can break that condition. Marginal coverage does not mean every site, backlog range or individual shipment has the same coverage. Slice coverage checks where the global margin is too small.
Separate width from usefulness
An interval that covers nearly every eventual clearance time by spanning several days may not help dispatch. Report empirical coverage, median or mean width and the proportion of intervals crossing the staffing decision boundary. Compare against a simple historical route interval on the same future cohort. The regression baseline anchors the point-prediction comparison; neither score alone establishes decision value.
Do not recycle the final test
Use calibration once under a frozen method, then evaluate interval behavior on a later test period. Changing alpha, strata or feature engineering after examining final-test misses turns the test into development data. Record training, calibration and test IDs, model version and prediction timestamp. The applied review checks all three roles.
Implementation
from math import ceil
# The predictor is already fitted; these are untouched calibration predictions.
calibration_actual = [6, 8, 5, 10, 7, 9, 4, 11, 8]
calibration_forecast = [5, 7, 6, 8, 7, 8, 6, 9, 7]
def conformal_margin(actual, predicted, alpha):
if not actual or len(actual) != len(predicted) or not 0 < alpha < 1:
raise ValueError("aligned calibration rows and valid alpha required")
residuals = sorted(abs(observed - forecast)
for observed, forecast in zip(actual, predicted))
rank = ceil((len(residuals) + 1) * (1 - alpha))
if rank > len(residuals):
raise ValueError("calibration sample too small for finite margin")
return residuals[rank - 1]
margin = conformal_margin(calibration_actual, calibration_forecast, alpha=0.20)
future_forecast = 8
future_interval = (future_forecast - margin, future_forecast + margin)
assert margin == 2
assert future_interval == (6, 10)Performance and operating cost
For N calibration cases, sorting residuals costs O(N log N) time and O(N) memory; serving an interval around an existing forecast is O(1). Recalibration requires new, eligible labels and a renewed exchangeability assessment, not merely a fast quantile calculation.
Common Mistakes
- Do not calculate the margin from training residuals or a tuned final test.
- Do not describe marginal coverage as a per-site or per-shipment guarantee.
- Do not substitute the largest residual when the requested rank exceeds the sample.
Read next
- Prediction interval coverage by operating slice
- Regression error slices and costly tails
- Group and time validation: split by the failure you expect in production
- Uncertainty and escalation review project
Continue the workflow: Prediction interval coverage by operating slice.
