Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Multilevel variance: separate within-group noise from new-group risk

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Random-intercept variation affects two observations from the same group and uncertainty for an unseen group.

Distinguish two variance components

Suppose depot handling time has a shared network mean, a depot-specific deviation and shift-level noise. The between-depot variance describes how far depot baselines spread; the within-depot variance describes differences among shifts inside a depot after modeled predictors. Two shifts in one depot share the depot deviation and are correlated even if their shift noise is independent. The intraclass correlation is the between component divided by the sum of both components. The cluster lesson shows why dependence changes uncertainty.

Predict known and new depots differently

For an observed depot, its shift history informs its deviation from the network mean and a partially pooled prediction can use that history. A newly opened depot has no measured deviation, so a prediction must include uncertainty from the between-depot distribution as well as future shift noise. Reusing an existing-depot interval for the new depot understates risk. The small program below calculates only a variance decomposition, not a predictive interval. The pooling lesson computes the conditional mean under a simple model.

Respect the grouping map

A driver may work across depots, while a depot may serve multiple routes. Those factors are crossed, not a neat chain of drivers nested in one depot. One random-intercept term cannot automatically absorb every source of correlation; specify actual assignment and repeated-measure structure. Check that groups are distinct and variance estimates are identifiable with the observed number of depots. A near-zero estimated between-depot component may reflect sparse data or a misspecified mean. Residual checks help find unmodeled patterns.

Tie prediction to the decision

Report whether the target is network mean, a known depot or a future depot. Show both variance components, depot and shift counts, and uncertainty for the chosen target. A fixed list of observed depot effects may describe current sites without supporting predictions for new sites; a random-effect model adds distributional assumptions. Hold the operational ranking if group composition changed during collection. The worked review adds a new depot and a mid-pilot equipment change.

Implementation

python
def depot_intraclass_correlation(between_depot_variance,
                                 within_depot_variance):
    if between_depot_variance < 0 or within_depot_variance <= 0:
        raise ValueError("nonnegative group and positive shift variance required")
    return between_depot_variance / (
        between_depot_variance + within_depot_variance)

assert depot_intraclass_correlation(9, 16) == 0.36
assert depot_intraclass_correlation(0, 16) == 0

Performance and operating cost

The ratio is O(1) time and space. Estimating components needs repeated work across shifts and groups; prediction for new groups needs an interval method. Ignoring group covariance makes a report cheap but can make its confidence intervals far too narrow.

Common Mistakes

  • Using a known-depot prediction interval for an unseen depot.
  • Treating many shifts as many independent depot assignments.
  • Forcing crossed driver and depot structure into one nested group.
  • Calling a small group variance proof that depots are identical.

Read next

Continue the workflow: Study heterogeneity: separate a mean effect from a new-site effect.

ai-data
applied-statistics
Storage details