A rolling-origin backtest trains on data available at each historical decision date and scores later cases while excluding outcomes that had not matured.
Rolling-origin retraining with an outcome embargo
Treat each origin like a real deployment
For a model used on Monday, only rows and labels available by Monday can train it. The next origin moves forward and repeats that decision. This tests the proposed retraining cadence against changing operations, unlike a random split that mixes seasons. Group and time validation also keeps repeat shipments and shared incidents together.
Embargo recent labels
If handoff outcomes need two days to settle, a Friday training snapshot cannot include Thursday labels even though the intake row exists. The code selects only cases whose label-ready timestamp precedes the origin and scores cases in the following window. The gap is a label-availability rule, not a cosmetic number of calendar days. Preserve source event time, arrival time and correction history. Delayed labels describes why maturity matters.
Refit the whole pipeline
Each origin must refit imputation, encoding, selection, calibration and the estimator using that origin’s eligible data. Reusing transformations fitted on the full table leaks future distribution information. A validation threshold chosen from all origins may overstate performance, so reserve a later untouched period for final comparison. Pipeline isolation covers the same boundary inside a fold.
Compare policies, not just scores
Evaluate a fixed model, a scheduled retrain and a drift-triggered retrain with the same future decision windows and costs. A new model may improve average error while increasing false negatives at one depot or expanding the manual-review queue. Report the cadence, training size, outcome maturity, calibration and rollback behavior. Review capacity is part of the policy.
Do not let late data rewrite the past
Backfills are useful for today’s training but were not visible to a historical decision. Backtesting from a current corrected table requires availability timestamps or archived snapshots. If those records are missing, state that the backtest is approximate. The release project requires a documented data clock before claiming a gain.
Implementation
# Intake day, label-ready day, alert outcome. Day numbers represent one archive.
cases = [(1, 3, 0), (2, 4, 1), (3, 6, 0), (4, 6, 1),
(5, 8, 0), (6, 9, 1), (7, 10, 1), (8, 12, 0),
(9, 12, 0), (10, 13, 1)]
origins = (6, 9)
forecast_horizon = 2
def rolling_windows(records, decision_days, horizon):
windows = []
for origin in decision_days:
train = [row for row in records if row[0] < origin and row[1] <= origin]
score = [row for row in records if origin <= row[0] < origin + horizon]
if not train or not score:
raise ValueError("origin lacks training or scoring rows")
windows.append((origin, train, score))
return windows
windows = rolling_windows(cases, origins, forecast_horizon)
assert [len(train) for _, train, _ in windows] == [4, 6]
assert [len(score) for _, _, score in windows] == [2, 2]Performance and operating cost
If each of R origins scans N archived rows, the simple backtest selector takes O(RN) time and O(N) memory per origin. Indexed timestamp selection can reduce repeated scans, while R complete model and preprocessing fits usually dominate compute. Archive availability timestamps if the result must support a release.
Common Mistakes
- Do not admit a label that became ready after the historical origin.
- Do not preprocess once on the full table before replaying origins.
- Do not use today’s corrected row as if it were visible in the past.
