A point forecast alone conceals uncertainty relevant to inventory. If a method produces an interval, state its intended coverage, construction version, horizon, and eligible backtest cases; then measure how often later actuals fell inside it. A prompt must not invent an '80% confidence' label from the width of a hand-drawn range. Coverage on a small history is descriptive and may shift after a process change. Examine whether misses concentrate in high-demand weeks, since average coverage can hide the cases most expensive for the planner.
Forecast prompts: label intervals and check coverage
Operational case
Aster's interval method had 20 eligible one-week backtest cases; 16 actual counts landed within their frozen intervals. Observed coverage is 16/20, or 80%, over that sample. For the next week it emits a point estimate of 54 and an interval from 45 to 65 kits. Depot stock capacity is 60, so the upper end crosses the current capacity. The memo states the interval's empirical history and the capacity exposure; it does not call the next week's outcome 80% certain or imply that 20 cases establish permanent calibration.
Backtest: 16 covered / 20 eligible = 80% observed coverage
Next-week point: 54 kits
Next-week interval: 45-65 kits
Current capacity: 60 kits
Upper interval exceeds capacity by 5 kitsPerformance and review cost
Coverage scoring across F frozen intervals is O(F) time. Calibration by horizon or demand band needs enough cases per slice; sparse slices should be labeled untested rather than assigned a misleading percentage. Wider intervals may improve coverage while making inventory decisions less specific. Report width and miss severity beside coverage, and keep the interval method version with the release packet.
Common Mistakes
- Do not call an arbitrary range a calibrated interval.
- Do not hide that only 20 cases support the observed coverage.
- Do not treat an interval boundary as a guaranteed maximum.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Chart prompts: review rendered and nonvisual outputs
- Aggregation prompts: pin the denominator and recompute the rate
- Forecast prompts: define target, horizon, and time grain
- Forecast prompts: enforce the as-of data boundary
- Forecast prompts: distinguish missing weeks from zero demand
- Forecast prompts: compare against a rolling baseline
- Forecast prompts: keep scenarios separate from predictions
- Forecast prompts: gate inventory recommendations on evidence
- Project: review Aster depot kit forecasts
- Forecast prompt decisions
