Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Survey weights: inspect concentration and the limits of effective sample size

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Unequal weights can make a nominally large survey depend heavily on a small number of respondents.

Inspect the final weights, not only their sum

After design, nonresponse and calibration adjustments, a late-shift respondent may represent many invited agents. Show minimum, median, upper tail, and total final weight, plus the share held by the largest few respondents or primary sampling units. If one branch dominates, a weighted overall satisfaction score can change sharply when a handful of answers move. The calibration lesson explains why these weights were created.

Use a diagnostic, not a replacement variance

The Kish-style unequal-weight effective count is the square of total weight divided by the sum of squared weights. It equals the raw sample count when all weights match and shrinks when a few weights dominate. The code computes that diagnostic. It does not capture clustering, stratification, outcome correlation with weights, or the finite population correction; therefore it is not a universal substitute for a design-based standard error. The PSU lesson keeps the actual sample design in the variance calculation.

Trim with a declared tradeoff

Capping extreme weights can lower variance but may reintroduce bias if the rare respondents truly represent underserved groups. Choose a cap by a documented rule, then compare estimates before and after trimming and, where appropriate, recalibrate to maintain known margins. A cap chosen only because it changes the headline in a favorable direction is not defensible. If an entire target subgroup never responded, trimming cannot solve its absent support.

Publish domain-level support

An overall estimate can have a reasonable effective count while an important late-shift or small-branch domain is supported by very few independent people. Report sample and weighted denominators by domain, response rates, weight concentration, and design-aware intervals. The project blocks a center ranking when a domain has no respondent support, even if the overall weighted score appears precise.

Implementation

python
def unequal_weight_effective_n(weights):
    if not weights or any(weight <= 0 for weight in weights):
        raise ValueError("positive weights required")
    return sum(weights) ** 2 / sum(weight ** 2 for weight in weights)

assert unequal_weight_effective_n([1, 1, 1, 1]) == 4
assert unequal_weight_effective_n([1, 1, 4]) == 2

Performance and operating cost

The diagnostic is O(n) time and O(1) extra space. Sorting weights for tail summaries costs O(n log n), and design-aware variance estimation may require replicate weights or PSU aggregation. A simple effective count is cheap but cannot capture every source of survey uncertainty.

Common Mistakes

  • Reporting raw respondent count as the precision of a highly weighted estimate.
  • Treating unequal-weight effective n as a complete design-effect calculation.
  • Trimming weights after seeing which cap makes a desired ranking appear.
  • Hiding a domain with no responding units behind a stable overall score.

Read next

ai-data
applied-statistics
Storage details