Build a daily receipt-volume forecast whose calendar, feature availability, baseline, backtest and release decision can all be inspected.
Project: forecast receipt volume with a reproducible rolling backtest
Create the dataset contract
Use daily receipt counts with a declared business time zone, complete-day flag and source snapshot. Include a confirmed zero day, an outage day and a late-arriving receipt. The outage must remain missing. Record the forecast origin and target horizon before selecting features. Calendar rules belong in the submitted manifest.
Establish the comparison
Implement last-value and seven-day seasonal naive baselines. If a candidate model is added, construct only features known at each origin and retain a two-day outcome delay when that matches the fixture. Run expanding or sliding rolling-origin backtests and explain the choice. Origin windows must not cross into their target periods.
Evaluate decisions
Report MAE, signed bias and record counts separately by horizon. Show errors for peak days and zero-demand days. If intervals are supplied, report empirical coverage and width on a later period; never tune them on that same period. A candidate should beat the baseline under a stated business cost, not merely a favorable pooled score.
Submit a release packet
Deliver the input fixture, calendar manifest, feature-availability table, forecast ledger, backtest results and a short decision record. Include tests that change a future value without affecting earlier features, remove a seasonal source day and simulate a stale actual. Show how a prior model version can be restored. Monitoring begins with that ledger.
Implementation
def release_gate(candidate_mae, baseline_mae, candidate_bias, max_abs_bias):
if baseline_mae <= 0 or candidate_mae < 0 or max_abs_bias < 0:
raise ValueError("invalid evaluation metrics")
return {"candidate_beats_baseline": candidate_mae < baseline_mae,
"bias_within_limit": abs(candidate_bias) <= max_abs_bias}Performance and operating cost
Generating O origins at H horizons creates O(OH) score records. A baseline is O(OH); fitted models add repeated training cost. Store one compact record per origin and horizon so an audit can reconstruct metrics without duplicating full datasets.
Common Mistakes
- Do not fill an outage day with zero demand.
- Do not select a model on the final evaluation period.
- Do not ship a model without a baseline and rollback record.
