A stratified estimate combines stratum rates using population shares, not the accidental share of completed responses.
Stratified survey estimates: weight toward the named population
Specify the survey target
A service desk surveys customers in metropolitan and rural regions after case closure. The target is the satisfaction rate among all eligible closed cases in a fixed month, not just people who responded. Define the eligibility frame, outcome wording, contact modes and known region totals. An oversampled rural group may be sensible for precision, but its raw sample share should not replace its population share in an overall estimate. The estimand and frame come first.
Construct the combined estimate
Estimate the satisfaction rate within each region, then multiply by that region’s share of eligible cases and sum. This simple poststratified calculation assumes respondents represent their region with respect to satisfaction; it cannot repair nonresponse differences within a region. If the sample had unequal inclusion probabilities inside regions, start from the documented design weights before any calibration. Base weights track that first selection step.
Expose missing cells and weight stress
A region with no valid respondents cannot contribute a measured regional rate. Do not silently redistribute its population share. Report invited, eligible, contacted and completed counts by region, plus missing satisfaction answers. Inspect the largest resulting weights and the effective sample size; a few high-weight responses can dominate the estimate. Weight diagnostics expose that concentration, while design variance addresses dependence that a weighted mean cannot describe.
Publish the scope of inference
Archive the population totals and date, response definition, region mapping, weighting formula and completed-cell counts. A weighted point estimate is still vulnerable to nonresponse associated with satisfaction after conditioning on region. Compare early and late responders, conduct a bounded nonresponse sensitivity analysis and describe unresolved selection risk. The project holds a statewide claim when one region has no usable response support.
Implementation
def population_weighted_satisfaction(responses_by_region, population_counts):
if set(responses_by_region) != set(population_counts):
raise ValueError("population and response regions differ")
total_population = sum(population_counts.values())
if total_population <= 0:
raise ValueError("population must be positive")
estimate = 0.0
for region, answers in responses_by_region.items():
if not answers or population_counts[region] < 0:
raise ValueError("empty or invalid region")
estimate += (population_counts[region] / total_population) * (sum(answers) / len(answers))
return estimate
estimate = population_weighted_satisfaction(
{"metro": [1, 0, 1, 1], "rural": [0, 1]},
{"metro": 700, "rural": 300})
assert round(estimate, 6) == 0.675
Performance and operating cost
The calculation is O(n + g) time and O(1) extra space beyond stored responses for n answers and g regions. Weight construction, contact attempts and nonresponse assessment dominate the real cost. The arithmetic yields a point estimate but does not supply a design-correct standard error or eliminate within-region selection bias.
Common Mistakes
- Using the completed-sample region share as the population share.
- Filling a no-response region with the overall respondent rate.
- Calling poststratification a cure for all nonresponse bias.
- Reporting the weighted estimate with an ordinary row-level standard error.
Read next
- Survey uncertainty: count sampled clusters, strata and weight concentration
- Project: audit a weighted contact-center satisfaction estimate
- Population, estimand and sampling frame: name the quantity before calculating
- Survey inclusion probabilities and base weights
- Weight extremes, effective sample size and trimming tradeoffs
Continue the workflow: Standardized rates: compare groups under one declared population mix.
Continue the workflow: Survey calibration: distinguish known joint cells from separate margins.
