A forecast should be compared with a simple baseline using the same origins, horizon, eligible rows, and error measure. Rolling-origin tests train or select information only through each origin, then score a later period. State whether the score is mean absolute error, signed error, or another measure and why that measure matches the decision. Keep per-origin errors, not just one average; a candidate that improves the average but fails during busy weeks may be unusable. The language model may explain results, while a deterministic evaluator calculates them.
Forecast prompts: compare against a rolling baseline
Operational case
Aster holds out four completed weeks with actual fulfilled counts 42, 47, 45, and 50. A prior-week baseline predicts 40, 42, 47, and 45; its absolute errors are 2, 5, 2, and 5, giving a mean absolute error of 3.5 kits. A candidate predicts 43, 46, 46, and 49; each absolute error is 1, giving mean absolute error 1.0. These fictional results support further review of the candidate, not a guarantee about the next week. Each prediction must have been frozen at its own origin.
Actual: 42, 47, 45, 50
Baseline: 40, 42, 47, 45 -> errors 2, 5, 2, 5 -> MAE 3.5
Candidate: 43, 46, 46, 49 -> errors 1, 1, 1, 1 -> MAE 1.0
Same four origins and one-week horizonPerformance and review cost
Scoring F frozen forecasts is O(F) time and O(1) memory when streamed. Training a new model for each origin can cost much more than scoring, so choose a realistic evaluation window and record compute cost. Do not compare a candidate trained on later revisions with a baseline generated from earlier snapshots. The clean baseline makes extra model complexity earn its place.
Common Mistakes
- Do not shuffle weeks across training and evaluation.
- Do not compare methods at different horizons or origins.
- Do not let the model calculate an unverified average from prose alone.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Evaluation sets: measure the failure cases that matter
- Numeric prompts: let code calculate and the model explain
- Forecast prompts: define target, horizon, and time grain
- Forecast prompts: enforce the as-of data boundary
- Forecast prompts: distinguish missing weeks from zero demand
- Forecast prompts: keep scenarios separate from predictions
- Forecast prompts: label intervals and check coverage
- Forecast prompts: gate inventory recommendations on evidence
- Project: review Aster depot kit forecasts
- Forecast prompt decisions
