A sample-size plan must account for the unit drawn, expected response, unequal weights and the smallest comparison the decision needs.
Precision budgets and effective sample size for weighted analyses
Plan backward from the decision
A team may need a defect-rate estimate for all support cases and a separate estimate for partner cases. The partner comparison often sets the review budget because it has fewer eligible units. Define an acceptable interval width and consequence of a wrong decision before choosing a count. A single rule such as “review 100 cases” ignores prevalence, clustering, response loss and the number of slices reported.
Inspect weight concentration
The common weight-only diagnostic (sum of weights)² / sum of squared weights measures how concentrated weights are. Four responses weighted 2, 2, 2 and 8 have a diagnostic of 196/76, about 2.58 rather than four equally weighted observations. That number is not a full design-based effective sample size when cases were drawn in clusters or have correlated outcomes. The sampling design remains primary.
Budget for response and review
If 47 completed reviews are needed and only about four in five selected cases are expected to yield an outcome, drawing exactly 47 is unlikely to suffice. Plan follow-up slots and record why outcomes fail. Use pilot response rates by stratum, then revise the allocation if a source proves harder to retrieve. Preserve a final audit sample separate from any active-learning queue used to improve a model.
Report uncertainty honestly
Compute intervals using the actual sampling units and weight scheme; a naive binomial interval can understate uncertainty after cluster selection or heavy weighting. Examine whether the interval crosses the operational threshold. Cluster resampling is one route when the design supports it, while sparse strata may need a wider interval or an explicit insufficient-data label.
Implementation
def weight_concentration_ess(weights):
if not weights or any(weight <= 0 for weight in weights):
raise ValueError("positive weights required")
total = sum(weights)
return total * total / sum(weight * weight for weight in weights)
assert round(weight_concentration_ess([2, 2, 2, 8]), 2) == 2.58Performance and operating cost
The diagnostic scans N weights in O(N) time and O(1) auxiliary space. Design-aware interval estimation may require cluster-level resampling, adding repeated O(N) passes and storage for group IDs.
Common Mistakes
- Do not treat a weight-only diagnostic as a complete variance calculation.
- Do not plan from completed reviews without response loss.
- Do not let an overall precision target hide a weak decision-critical stratum.
Read next
- Stratified sampling and design weights for operational estimates
- Nonresponse bias: diagnose the missing outcomes before adjusting
- Project: estimate support-case defects from an auditable sample
- Standard error and cluster bootstrap: resample the independent unit
Continue the workflow: Sensitivity analysis: find which assumptions can reverse a decision.
Continue the workflow: Decision objectives: define the action before optimizing a score.
Continue the workflow: Prior predictive checks for a binary service rate.
Continue the workflow: Weight extremes, effective sample size and trimming tradeoffs.
Continue the workflow: Simulation error and replication budgets.
