Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Regression uncertainty: inspect changing residual spread and high influence

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Heteroskedasticity changes error variance across observations; influence determines how strongly each row can shape a fitted regression.

Define the coefficient before its standard error

A carrier analyst regresses route time on distance, parcel load, depot and schedule. The slope may describe a conditional association or a policy effect only under different design assumptions. Plot residuals against fitted time, distance, and depot before trusting the usual constant-variance standard error. A widening residual cloud suggests that long routes have different noise from short ones. The estimand lesson defines what the coefficient can represent; residual diagnostics expose misspecification beyond changing variance.

Separate a covariance repair from a model repair

An HC3-style heteroskedasticity-consistent covariance uses squared residuals scaled by influence, roughly dividing each residual by one minus its hat value before building the covariance matrix. The code calculates that ingredient, not a complete coefficient covariance. A variance-aware covariance can alter the uncertainty attached to the same fitted coefficients when errors are heteroskedastic, but it does not remove confounding, missing nonlinear structure, or dependence between repeated routes from one depot. A row with influence nearly one deserves direct inspection because its adjusted residual can dominate.

Choose the independent unit

If many routes share a driver, depot, or dispatch day, errors may move together. HC3 alone addresses changing row-level variance under independence assumptions; it does not create independent observations where clusters exist. A cluster-aware covariance or cluster bootstrap needs enough independent groups and the right grouping level. The cluster lesson addresses that boundary. Do not select a standard-error formula solely because it makes a desired slope significant.

Report sensitivity on the same effect scale

Show the coefficient and its unit, ordinary and justified alternative standard errors, influence distribution, influential rows, and the operational consequences of excluding any records. If a few high-influence routes determine the direction, the model may be extrapolating beyond common route support. The project requires a frozen target population and a depot-level dependence review before a route-time claim is released.

Implementation

python
def hc3_residual_factors(residuals, leverage_values):
    if not residuals or len(residuals) != len(leverage_values):
        raise ValueError("aligned residuals and influence required")
    factors = []
    for residual, influence in zip(residuals, leverage_values):
        if not 0 <= influence < 1:
            raise ValueError("influence must be in [0, 1)")
        factors.append((residual / (1 - influence)) ** 2)
    return factors

factors = hc3_residual_factors([3.0, 3.0], [0.2, 0.8])
assert factors[1] > 10 * factors[0]

Performance and operating cost

Computing n HC3 residual factors is O(n) time and output space. Full covariance construction also uses the design matrix and matrix operations. Refitting with a better functional form or clustering strategy costs more, but multiplying residuals alone cannot fix those model boundaries.

Common Mistakes

  • Calling HC3 a cure for confounding or a wrong mean model.
  • Using row-independent HC3 when many observations share one depot shock.
  • Deleting a high-influence route only because it weakens the preferred result.
  • Treating a slope measured in minutes per kilometer as a unitless effect.

Read next

ai-data
applied-statistics
Storage details