Compare random-row and geographically separated validation, then qualify the forecast for the districts where it will be used.
Project: audit branch-demand predictions for new districts
Define the deployment map
A retailer has demand histories for existing branches and plans inventory for districts with no operating branch. Freeze the target period, eligible sites, coordinate system, district boundaries and whether the claim is for new districts or another month in the same district. Check duplicate branch IDs, locations at district borders and any data source whose feature was computed after the forecast date. A random train-test split answers a much easier question than this new-district decision.
Inspect local similarity
Build a defensible neighbor graph from travel distance or another business distance measure and inspect demand or residual similarity. The graph need not be complex, but its rule must be recorded before the statistic is interpreted. The spatial diagnostic makes shared local structure visible. Demand can also move with season, so do not mistake later outcomes from nearby branches for features known at prediction time.
Hold out districts with a gap
Set multiple test regions, place a buffer around each and fit the entire preprocessing and model pipeline using training regions only. Compare error, calibration and high-demand miss rate by held-out district against a random-row baseline, with sample counts for both. The split lesson provides the minimum geography guard; the time lesson provides the date guard. The geographic estimate may have larger error, which is useful evidence rather than a reason to hide it.
Publish a bounded recommendation
The packet includes a map of train, gap and test sites, coordinate quality, fold assignment, regional sample sizes, feature availability, model version and separate metrics for familiar versus new districts. The gate below blocks a launch score if geographic separation or train-only preprocessing was skipped. Passing it allows planning review, not a claim that every unsampled district behaves like the held-out districts.
Implementation
def spatial_forecast_gate(audit):
if not audit["deployment_geography_defined"]:
return "hold:target-geography"
if not audit["buffered_regions_tested"]:
return "hold:spatial-validation"
if not audit["preprocessing_fit_on_train_only"]:
return "hold:feature-leakage"
if not audit["regional_errors_reported"]:
return "hold:regional-error"
return "review:new-district-forecast"
audit = {"deployment_geography_defined": True, "buffered_regions_tested": True,
"preprocessing_fit_on_train_only": False, "regional_errors_reported": True}
assert spatial_forecast_gate(audit) == "hold:feature-leakage"
assert spatial_forecast_gate({**audit, "preprocessing_fit_on_train_only": True}) == "review:new-district-forecast"
Performance and operating cost
The gate is O(1). Geographic partitioning takes O(n) per held-out region; repeated fitting can be much more expensive. Training on fewer sites is an intentional cost of asking whether a model can work away from nearby observations.
Common Mistakes
- Hiding geographic test error behind a random-row score.
- Allowing held-out district outcomes into target encodings or preprocessing.
- Using unprojected coordinates as kilometers.
- Claiming the held-out districts represent every future market.
Read next
- Spatial dependence: inspect neighbor similarity before inference
- Buffered spatial holdout: test prediction away from nearby training sites
- Rolling-origin backtests: make every forecast from information available then
- Clustered regression: count independent groups before trusting precision
- Population, estimand and sampling frame: name the quantity before calculating
