A survey’s standard error must reflect how units were selected; the respondent count alone does not reveal independent information.
Survey uncertainty: count sampled clusters, strata and weight concentration
Map the actual selection design
The service desk samples contact-center branches, then cases inside each branch. Cases from one branch share staffing, customer mix and workflow, so they are not independent draws from the full population. Record each sampled branch as a primary sampling unit, each design stratum, inclusion and response-adjusted weights, and the analysis population. Cluster resampling explains why replacing individual cases can make uncertainty too small.
Check whether variance is identifiable
A stratum represented by one sampled branch lacks direct within-stratum between-branch variation for a simple design-based variance calculation. A chosen lonely-cluster method or pooling rule requires a documented design decision; setting variance to zero is not acceptable. Count distinct branches per stratum before fitting a standard error. The diagnostic code below only checks support and weight concentration; it intentionally does not pretend to be a full survey-variance estimator.
Separate two information limits
Kish effective sample size, the square of total weight divided by the sum of squared weights, describes weight concentration. It does not account for clustering and is not an adjusted number of independent branches. A large respondent count can coexist with few sampled branches and extreme weights. Use software and variance methods that encode the original strata, primary units and weights for reportable confidence intervals. Precision budgeting helps set collection targets, while coverage explains the repeated-sample claim.
Report design sensitivity
Compare weighted and unweighted estimates as a diagnostic, not as competing truths. Show completion by stratum, sampled branches, weight range, effective count and interval method. If a region or stratum is poorly supported, mark it as limited rather than forcing a narrow statewide interval. Population weighting provides the point estimate; the audit project blocks publication when a one-branch stratum cannot support the promised precision.
Implementation
def survey_design_check(records):
if not records:
raise ValueError("survey records required")
branches = {}
weights = []
for record in records:
if record["weight"] <= 0:
raise ValueError("weights must be positive")
branches.setdefault(record["stratum"], set()).add(record["branch_id"])
weights.append(record["weight"])
effective_count = sum(weights) ** 2 / sum(weight ** 2 for weight in weights)
return {"branches_per_stratum": {key: len(value) for key, value in branches.items()},
"weight_effective_count": effective_count,
"supported": all(len(value) >= 2 for value in branches.values())}
survey = [{"stratum": "metro", "branch_id": "m-47", "weight": 2},
{"stratum": "metro", "branch_id": "m-82", "weight": 3},
{"stratum": "rural", "branch_id": "r-31", "weight": 7}]
report = survey_design_check(survey)
assert not report["supported"]
assert report["weight_effective_count"] < len(survey)
Performance and operating cost
The support scan is O(n) expected time and O(p) space for n responses and p distinct primary units. A design-aware variance calculation costs more and depends on the survey plan. The effective count is a warning about variable weights, not a shortcut to a confidence interval that ignores clustering.
Common Mistakes
- Treating every case inside one sampled branch as independent.
- Using weight effective count as if it accounted for cluster dependence.
- Assigning zero variance to a stratum with one sampled branch.
- Discarding strata and primary-unit IDs after calculating a weighted mean.
Read next
- Stratified survey estimates: weight toward the named population
- Project: audit a weighted contact-center satisfaction estimate
- Standard error and cluster bootstrap: resample the independent unit
- Confidence intervals: interpret coverage and precision honestly
- Precision budgets and effective sample size for weighted analyses
Continue the workflow: Survey weights: inspect concentration and the limits of effective sample size.
Continue the workflow: Finite-frame sampling: adjust uncertainty for sampling without replacement.
Continue the workflow: Ratio uncertainty: retain numerator-denominator covariance.
