Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Clustered regression: count independent groups before trusting precision

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Cluster-aware uncertainty allows residuals within a group to co-move, while identification and finite-group limits remain separate questions.

Map the shared shock

A route-time policy changes at depot level, yet an extract has tens of thousands of parcel rows. Parcels from the same depot-day share staff, weather, loading delays, and policy assignment. Treating parcel rows as independent can give a tiny standard error without many independent policy comparisons. The correct group may be depot, depot-day, or another unit depending on assignment and shock persistence; it must be justified from the data-generating process. The randomization-unit lesson is the first audit for a trial.

Aggregate score contributions within group

A cluster-variance-aware sandwich covariance builds group-level sums of each row’s design-vector contribution times its residual, then combines outer products of those group sums. The code below computes one scalar predictor score total per group to make the aggregation visible. It is not a complete cluster covariance or a small-cluster correction. If a shock persists across days at one depot, grouping only by depot-day may still split a dependent process. HC3 treats changing variance but cannot substitute for that dependence map.

Check the number and balance of groups

Thirty thousand rows in four depots do not give thirty thousand independent assignment units. With few clusters or one cluster dominating treatment exposure, asymptotic standard errors can be unreliable. Inspect cluster counts by arm, size distribution, treatment support, and whether a cluster ever changes treatment status. A cluster bootstrap should resample whole independent groups, not parcels; with very few groups it also has limits. The bootstrap lesson explains the resampling unit.

Keep inference and bias separate

A cluster-aware interval can widen uncertainty but cannot undo an observational policy rollout confounded by depot readiness. Report the coefficient with its scale, clustering unit and rationale, independent group count, and sensitivity to plausible group definitions. The project blocks a precise claim when assignment groups and inference groups do not align. A larger data extract is not a replacement for new independent depots.

Implementation

python
from collections import defaultdict

def predictor_residual_scores(route_rows):
    totals = defaultdict(float)
    for depot_id, predictor_value, residual in route_rows:
        if not depot_id:
            raise ValueError("depot ID required")
        totals[depot_id] += predictor_value * residual
    return dict(totals)

scores = predictor_residual_scores([("east", 2, 3), ("east", 4, -1),
                                    ("west", 3, 2)])
assert scores == {"east": 2.0, "west": 6.0}

Performance and operating cost

The group-score accumulation is O(n) time and O(g) space for n rows and g groups. A full covariance also needs matrix products and a defensible finite-group procedure. The expensive step is finding the real dependence boundary, not summing another million parcel rows.

Common Mistakes

  • Clustering by a smaller unit than the assignment or persistent shock.
  • Calling thousands of parcels thousands of independent policy units.
  • Assuming a cluster correction fixes confounded policy assignment.
  • Using an asymptotic interval with only a handful of uneven groups without qualification.

Read next

Continue the workflow: Spatial dependence: inspect neighbor similarity before inference.

ai-data
applied-statistics
Storage details