Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Rolling-origin backtests: make every forecast from information available then

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A rolling-origin evaluation moves the forecast date forward and uses only data available before each origin to predict a declared future horizon.

Freeze the production timing

A service center forecasts next week’s incoming cases every Friday. The input snapshot at a Friday origin must contain only cases, holidays and staffing plans known by that Friday; later corrections and post-week totals cannot leak backward. For a four-week horizon, the answer is the demand four weeks out, not the next observed week. Choose train length, update cadence, horizon and final audit period before model comparison. The serial-dependence lesson explains why randomly shuffled rows give misleading certainty.

Generate comparable origins

At each origin, fit or update the candidate using observations ending before a declared gap, then predict the target horizon. The gap may reflect reporting delay or operational embargo. The code returns simple index windows for an expanding history; it does not fit a model. Every forecast must be stored with its origin, target period, horizon and model version. If a target period lacks a final count, flag it as pending rather than treating it as zero.

Compare against a plain baseline

A seasonal last-year value or recent-week average can be a serious baseline. Use the same origins and target periods for every candidate, then score both overall and by horizon, region and known peak periods. Aggregate errors from many adjacent origins can be correlated because their training sets and forecast targets overlap. Avoid calling the number of forecasts the number of independent policy trials. The interval lesson checks whether uncertainty grows appropriately with horizon.

Preserve a final untouched window

Tune models on earlier backtest origins, reserve later origins for final comparison, and then monitor live forecasts using their original archived versions. Rebuilding a forecast after the target week using newly corrected history is useful for diagnosis but not a valid historical prediction. The project requires timestamped feature snapshots and a later-period score before staffing changes are justified.

Implementation

python
def expanding_forecast_windows(total_periods, minimum_train,
                               horizon, gap=0):
    if minimum_train <= 0 or horizon <= 0 or gap < 0 or        total_periods < minimum_train + gap + horizon:
        raise ValueError("history cannot support the requested windows")
    return [(tuple(range(0, train_end)), train_end + gap + horizon - 1)
            for train_end in range(minimum_train,
                                   total_periods - gap - horizon + 1)]

windows = expanding_forecast_windows(12, 5, 3, gap=1)
assert windows[0] == ((0, 1, 2, 3, 4), 8)
assert windows[-1][1] == 11

Performance and operating cost

The window generator creates O(k n) stored indices across k origins and up to n training periods; streaming slices avoids the list cost. Refitting every origin is more expensive but reproduces the production forecast. Random train-test shuffling is faster and answers the wrong timing question.

Common Mistakes

  • Using a feature corrected after the forecast origin.
  • Scoring a four-week forecast against next week.
  • Treating pending target counts as zero.
  • Tuning and reporting accuracy on the same final origins.

Read next

Continue the workflow: Buffered spatial holdout: test prediction away from nearby training sites.

ai-data
applied-statistics
Storage details