Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Rolling-origin backtests: rehearse the forecast as it would have run

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A backtest evaluates forecasts at successive historical origins using only data available before each origin.

Move origin forward

Choose an initial training window, issue a forecast, score its target period, then move the origin. Expanding windows retain all prior history; sliding windows discard older periods. State which behavior is intended and why. A random row split lets future observations train a model for earlier predictions and invalidates the deployment simulation.

Leave a data gap

If outcomes finalize two days after submission, the training window at today’s origin must stop before those unresolved days. A gap also protects against lookahead from lagged features or source processing delays, but its size should come from availability facts. Feature lineage must be reconstructed for every origin.

Avoid overlapping-score confusion

A seven-day forecast issued daily creates overlapping target windows. Scoring every daily origin is still useful, yet errors are correlated and naive confidence calculations can overstate independent evidence. Report origin count, horizon and evaluation period. Keep one final recent holdout untouched while choosing features and model settings.

Test the index arithmetic

With 35 daily observations, an initial 14-day training window, a two-day gap and a three-day forecast horizon, the first target begins at index 16. Its target indices are 16, 17 and 18. Confirm later origins never train on their own target rows and stop before the horizon exceeds available history.

Implementation

python
def rolling_origins(observation_count, min_train=14, gap=2, horizon=3):
    if min_train < 1 or gap < 0 or horizon < 1:
        raise ValueError("invalid backtest window")
    for train_end in range(min_train, observation_count - gap - horizon + 1):
        target_start = train_end + gap
        yield range(0, train_end), range(target_start, target_start + horizon)

Performance and operating cost

For O origins, generating index windows costs O(O) iterator overhead, but refitting a model can dominate at O times its training cost. Store compact origin and score records rather than copying every training slice unless reproducibility requires it.

Common Mistakes

  • Do not random-split a temporally ordered forecast task.
  • Do not choose a gap by habit when source availability has a different delay.
  • Do not treat overlapping forecast errors as independent replicates.

Read next

ai-data
time-series
Storage details