Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Split conformal intervals for clearance forecasts

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Split conformal regression uses absolute errors from a held-out calibration set to place an empirical margin around a fixed point predictor.

Reserve a calibration cohort

A shipment model predicts hours until clearance. Fit the predictor on training shipments, select its settings on development data, then freeze it before scoring a separate calibration cohort. None of those calibration outcomes may have influenced the model or preprocessing. The margin comes from absolute calibration residuals, not training residuals. The split guide explains why related shipments and future timestamps must not leak between roles.

Choose a finite-sample rank

For a requested miscoverage fraction alpha and n calibration residuals, use the rank ceiling of (n plus one) times (one minus alpha), with one-based indexing. When that rank exceeds n, a finite margin cannot support the requested finite-sample claim without a wider convention; the code raises instead of quietly taking the largest observed error. The resulting interval is the point forecast plus or minus the selected residual. It is a teaching implementation of the split procedure.

State the coverage scope

The familiar marginal coverage guarantee relies on calibration and future examples being exchangeable under the prediction procedure. A later warehouse policy, selective labeling or dependent repeated scans can break that condition. Marginal coverage does not mean every site, backlog range or individual shipment has the same coverage. Slice coverage checks where the global margin is too small.

Separate width from usefulness

An interval that covers nearly every eventual clearance time by spanning several days may not help dispatch. Report empirical coverage, median or mean width and the proportion of intervals crossing the staffing decision boundary. Compare against a simple historical route interval on the same future cohort. The regression baseline anchors the point-prediction comparison; neither score alone establishes decision value.

Do not recycle the final test

Use calibration once under a frozen method, then evaluate interval behavior on a later test period. Changing alpha, strata or feature engineering after examining final-test misses turns the test into development data. Record training, calibration and test IDs, model version and prediction timestamp. The applied review checks all three roles.

Implementation

python
from math import ceil

# The predictor is already fitted; these are untouched calibration predictions.
calibration_actual = [6, 8, 5, 10, 7, 9, 4, 11, 8]
calibration_forecast = [5, 7, 6, 8, 7, 8, 6, 9, 7]

def conformal_margin(actual, predicted, alpha):
    if not actual or len(actual) != len(predicted) or not 0 < alpha < 1:
        raise ValueError("aligned calibration rows and valid alpha required")
    residuals = sorted(abs(observed - forecast)
                       for observed, forecast in zip(actual, predicted))
    rank = ceil((len(residuals) + 1) * (1 - alpha))
    if rank > len(residuals):
        raise ValueError("calibration sample too small for finite margin")
    return residuals[rank - 1]

margin = conformal_margin(calibration_actual, calibration_forecast, alpha=0.20)
future_forecast = 8
future_interval = (future_forecast - margin, future_forecast + margin)
assert margin == 2
assert future_interval == (6, 10)

Performance and operating cost

For N calibration cases, sorting residuals costs O(N log N) time and O(N) memory; serving an interval around an existing forecast is O(1). Recalibration requires new, eligible labels and a renewed exchangeability assessment, not merely a fast quantile calculation.

Common Mistakes

  • Do not calculate the margin from training residuals or a tuned final test.
  • Do not describe marginal coverage as a per-site or per-shipment guarantee.
  • Do not substitute the largest residual when the requested rank exceeds the sample.

Read next

Continue the workflow: Prediction interval coverage by operating slice.

ai-data
machine-learning
Storage details