Unequal weights can make a nominally large survey depend heavily on a small number of respondents.
Survey weights: inspect concentration and the limits of effective sample size
Inspect the final weights, not only their sum
After design, nonresponse and calibration adjustments, a late-shift respondent may represent many invited agents. Show minimum, median, upper tail, and total final weight, plus the share held by the largest few respondents or primary sampling units. If one branch dominates, a weighted overall satisfaction score can change sharply when a handful of answers move. The calibration lesson explains why these weights were created.
Use a diagnostic, not a replacement variance
The Kish-style unequal-weight effective count is the square of total weight divided by the sum of squared weights. It equals the raw sample count when all weights match and shrinks when a few weights dominate. The code computes that diagnostic. It does not capture clustering, stratification, outcome correlation with weights, or the finite population correction; therefore it is not a universal substitute for a design-based standard error. The PSU lesson keeps the actual sample design in the variance calculation.
Trim with a declared tradeoff
Capping extreme weights can lower variance but may reintroduce bias if the rare respondents truly represent underserved groups. Choose a cap by a documented rule, then compare estimates before and after trimming and, where appropriate, recalibrate to maintain known margins. A cap chosen only because it changes the headline in a favorable direction is not defensible. If an entire target subgroup never responded, trimming cannot solve its absent support.
Publish domain-level support
An overall estimate can have a reasonable effective count while an important late-shift or small-branch domain is supported by very few independent people. Report sample and weighted denominators by domain, response rates, weight concentration, and design-aware intervals. The project blocks a center ranking when a domain has no respondent support, even if the overall weighted score appears precise.
Implementation
def unequal_weight_effective_n(weights):
if not weights or any(weight <= 0 for weight in weights):
raise ValueError("positive weights required")
return sum(weights) ** 2 / sum(weight ** 2 for weight in weights)
assert unequal_weight_effective_n([1, 1, 1, 1]) == 4
assert unequal_weight_effective_n([1, 1, 4]) == 2
Performance and operating cost
The diagnostic is O(n) time and O(1) extra space. Sorting weights for tail summaries costs O(n log n), and design-aware variance estimation may require replicate weights or PSU aggregation. A simple effective count is cheap but cannot capture every source of survey uncertainty.
Common Mistakes
- Reporting raw respondent count as the precision of a highly weighted estimate.
- Treating unequal-weight effective n as a complete design-effect calculation.
- Trimming weights after seeing which cap makes a desired ranking appear.
- Hiding a domain with no responding units behind a stable overall score.
Read next
- Survey calibration: distinguish known joint cells from separate margins
- Project: calibrate a contact-center survey without hiding sparse shifts
- Survey uncertainty: count sampled clusters, strata and weight concentration
- Stratified survey estimates: weight toward the named population
- Standard error and cluster bootstrap: resample the independent unit
