Review a fictional one-week replacement-kit forecast for Aster Repair depot D-47. Start by defining fulfilled kits as the target, a depot-local Monday-to-Monday week as the grain, and Sunday 18:00 as the forecast origin. A partial current week is not a completed observation. The task is to prepare a human-review packet; no purchase or transfer is authorized by this exercise.
Project: review Aster depot kit forecasts
Rebuild an honest backtest
Use origin-specific snapshots. A correction posted Wednesday and a promotion approved Tuesday were unavailable at the earlier Sunday origin. Mark a scanner-export gap as unknown, a scheduled-closure week as an observed zero, and a 32-kit stockout week as censored fulfilled demand. Four held-out actuals are 42, 47, 45, and 50. The prior-week baseline predicts 40, 42, 47, and 45, producing mean absolute error 3.5. A candidate predicts 43, 46, 46, and 49, producing mean absolute error 1.0. Both methods must use the same one-week origins and eligible periods.
Keep scenarios and decisions separate
The candidate's next-week point forecast is 54 kits. A promotion scenario adds eight only if that unapproved event occurs, yielding 62; it is not the baseline. Sixteen of 20 frozen interval backtests covered the actual, an observed 80% for that small set. The current interval is 45 to 65 kits, while storage capacity is 60. Ask for usable stock, inbound receipt timing, and lead time before proposing an order quantity. The decision memo should show the five-kit interval exposure above capacity without treating the interval as a hard maximum.
Baseline backtest errors: 2, 5, 2, 5 -> MAE 3.5 kits
Candidate backtest errors: 1, 1, 1, 1 -> MAE 1.0 kit
Point: 54 kits; conditional promotion: 62 kits
Interval: 45-65 kits; observed coverage: 16/20
Capacity: 60 kits; usable stock and inbound timing: unknown
Order quantity: pending reviewPerformance and review cost
Aggregate N events in O(N) expected time, score F frozen forecasts in O(F), and check interval coverage in O(F). Model fitting and data reconstruction may dominate those arithmetic checks. The review packet must include the as-of snapshot, rejected periods, target calendar, baseline and candidate errors, interval method, scenario approvals, and decision owner. A fluent forecast summary without those records is not reproducible.
Common Mistakes
- Do not use a partial week or a future correction as past data.
- Do not fill an ingest gap with zero.
- Do not present the conditional 62-kit scenario as the 54-kit baseline.
- Do not order from the point forecast while stock and lead time are unknown.
Related lessons
- Forecast prompts: define target, horizon, and time grain
- Forecast prompts: enforce the as-of data boundary
- Forecast prompts: distinguish missing weeks from zero demand
- Forecast prompts: compare against a rolling baseline
- Forecast prompts: keep scenarios separate from predictions
- Forecast prompts: label intervals and check coverage
- Forecast prompts: gate inventory recommendations on evidence
- Numeric prompts: let code calculate and the model explain
- Evaluation leakage: keep the release test independent
- Forecast prompt decisions
