Between-study variation widens uncertainty for a future setting even when the average effect is estimated precisely.
Study heterogeneity: separate a mean effect from a new-site effect
Distinguish two sources of spread
Warehouse pilots differ because of finite samples and because site operations may produce genuinely different route effects. Within-study variance quantifies uncertainty in each estimate; between-study variance describes dispersion of underlying site effects under a random-effects model. A precise average across several sites can coexist with a wide range of plausible effects at the next site. The code adds mean-estimation variance and between-site variance to show that difference; it does not calculate a calibrated prediction interval. Effect alignment comes first.
Do not let one number replace inspection
A heterogeneity percentage or a fitted between-site variance is sensitive to the number and precision of available studies. With few sites, its estimate is unstable and an apparently zero value does not establish identical effects. Inspect each site result, route mix, equipment version, assignment method, and plausible effect modifiers. If the intervention was implemented differently, an average may not represent a coherent policy. The multilevel lesson explains the analogous distinction between known and unseen groups.
Keep the target of inference explicit
A confidence interval for the average underlying effect addresses the mean over the modeled set of sites. A prediction interval for a new comparable site must also allow between-site variation and uncertainty in its estimate, with careful treatment when the study count is small. A narrow average interval does not promise that every future deployment saves time. If the next warehouse has a different shipment mix, even the random-effects prediction may not transport. Standardization can address measured mix changes, not unknown implementation differences.
Tell the reviewer what changed
Report common-effect and heterogeneity-aware summaries as sensitivity analyses, not as a contest where the more favorable answer wins. Show site estimates, study count, between-site variance, mean interval, new-site interval when defensible, and reasons for any excluded report. The operational decision may require a limited local trial if the new-site range includes harm. The synthesis project makes that release gate concrete.
Implementation
from math import sqrt
def new_site_standard_deviation(mean_estimate_variance,
between_site_variance):
if mean_estimate_variance < 0 or between_site_variance < 0:
raise ValueError("variances cannot be negative")
return sqrt(mean_estimate_variance + between_site_variance)
assert new_site_standard_deviation(0.25, 4.0) > 2
assert new_site_standard_deviation(0.25, 0.0) == 0.5
Performance and operating cost
Combining two supplied variance components is O(1). Estimating between-site variance, checking compatibility, and calibrating a small-study prediction interval are the expensive steps. A simple average interval is cheaper but answers a narrower question than deployment risk.
Common Mistakes
- Calling the confidence interval for the mean a range for every future site.
- Declaring effects identical because fitted heterogeneity is zero in a tiny set.
- Selecting a synthesis model after seeing which answer supports rollout.
- Transporting a site range to a warehouse with a different implementation.
Read next
- Combining studies: align effects, units, and independent evidence
- Project: combine warehouse pilots without double-counting evidence
- Multilevel variance: separate within-group noise from new-group risk
- Standardized rates: compare groups under one declared population mix
- Confidence intervals: interpret coverage and precision honestly
