Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Spatial dependence: inspect neighbor similarity before inference

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Spatial autocorrelation means nearby locations can have related values, reducing the information gained from treating every location as an independent replicate.

Define neighbors before seeing outcomes

A chain of retail branches shares catchment areas, weather and local promotions. A demand spike at one branch may coincide with a spike nearby. Represent spatial proximity with a documented neighbor graph or distance rule before looking for a favorable pattern. The code calculates a simple global Moran statistic from a symmetric neighbor graph; it uses each listed directed edge in the weight sum. It is a descriptive diagnostic. Its value depends on the graph, geographic extent and outcome distribution, so a positive number alone is not a calibrated significance test.

Read the statistic with its weight map

The statistic compares cross-products of centered values on neighboring branches with total centered variation. Positive values suggest neighboring branches tend to be similar under the chosen graph; negative values suggest opposite-valued neighbors. An isolated branch has no neighbor contribution but still affects the global mean and denominator. Changing from road travel time to straight-line distance can change the result. Report the weights, islands and number of edges. Serial dependence is the time analogue, but spatial effects need geographic design.

Separate dependence from explanation

Nearby high-demand branches may reflect shared population density, a regional campaign, a data-entry convention or a real spillover. A spatial pattern does not identify which mechanism operates. Assess residual spatial structure after a justified model, and calibrate a reference distribution only under a null and permutation scheme that respects the sampling and spatial support. If branch locations were selected by business opportunity, permuting values as though they were exchangeable can overstate certainty.

Change the validation question

A random row split often puts neighboring branches in training and test sets, making prediction at nearby sites look easier than prediction in a new region. Use a holdout design that matches the intended deployment geography. The buffered holdout lesson addresses near-neighbor leakage; the project checks the operational claim. Spatial dependence is not a request to put latitude and longitude into every regression regardless of purpose.

Implementation

python
def moran_neighbor_statistic(values_by_branch, neighbors):
    branch_ids = set(values_by_branch)
    if len(branch_ids) < 3 or set(neighbors) != branch_ids:
        raise ValueError("at least three aligned branches required")
    center = sum(values_by_branch.values()) / len(branch_ids)
    deviations = {branch: value - center for branch, value in values_by_branch.items()}
    denominator = sum(value * value for value in deviations.values())
    edge_total = weighted_sum = 0
    for branch, adjacent in neighbors.items():
        if any(other not in branch_ids or other == branch or
               branch not in neighbors[other] for other in adjacent):
            raise ValueError("symmetric, non-self neighbor graph required")
        for other in adjacent:
            weighted_sum += deviations[branch] * deviations[other]
            edge_total += 1
    if not denominator or not edge_total:
        raise ValueError("variation and at least one edge required")
    return len(branch_ids) * weighted_sum / (edge_total * denominator)

branch_values = {"A": 1, "B": 1, "C": 9, "D": 9}
branch_neighbors = {"A": {"B"}, "B": {"A", "C"},
                    "C": {"B", "D"}, "D": {"C"}}
assert round(moran_neighbor_statistic(branch_values, branch_neighbors), 3) == 0.333

Performance and operating cost

With n branches and e directed neighbor entries, this diagnostic takes O(n+e) time and O(n) extra space. Permutation calibration repeats the calculation many times and depends on a valid null. A larger neighbor matrix is not a substitute for specifying the deployment geography.

Common Mistakes

  • Reading the sign of one statistic as proof of a regional cause.
  • Changing the neighbor definition until a desired result appears.
  • Assuming a random branch split tests new-region performance.
  • Treating adjacent branches as independent even after a residual spatial pattern remains.

Read next

ai-data
applied-statistics
Storage details