Build a traceable case-level survey estimate from inclusion probabilities through response adjustment, calibration and precision checks.
Project: audit weights for a support-handoff survey
Freeze the population and draw
Create a 1,200-case frame with 800 standard and 400 partner handoffs. Randomly invite 80 cases from each queue, retaining draw IDs and known selection probabilities. The survey item asks whether the next step after the handoff was mostly or completely clear. The measurement contract defines the case and the recall window; weights cannot repair a different question.
Reconcile respondents and base weights
Use 48 standard and 32 partner completed surveys. Verify that every completion belongs to one invitation and that invitation IDs are unique. Base weights are 10 and five, reconstructing the frame from invitations. Response-adjusted weights are 800/48 and 400/32, reconstructing the frame from respondents. Record both stages rather than overwriting one weight column. The design lesson explains the distinction.
Calculate an item estimate with denominators
Suppose 28 of 48 standard respondents and 12 of 32 partner respondents choose the top two clarity categories. The unweighted respondent share is 40/80, or 50%. With queue-adjusted weights, the weighted share is about 51.4% across frame cases under the within-queue response assumption. Keep any item skips separate; this fixture assumes all 80 completed respondents provided a substantive answer. Response adjustment does not guarantee that assumption is true.
Test calibration and concentration
Add known frame channel totals of 720 email and 480 phone cases. Compare those with weighted respondent channel totals, then rake to both queue and channel controls if the respondent data support every required category. Check residuals and investigate large adjustment factors. Calculate the weights-only effective size and subgroup respondent counts before and after calibration. Precision diagnostics reveal what a final percentage hides.
Publish the audit packet
Save frame version, inclusion probabilities, invitation and response counts, item codebook, base and adjusted weights, control totals, calibration tolerance, weight distribution, weighted and unweighted estimates, and a design-aware uncertainty plan. Compare respondents and nonrespondents on known case features. If a cell lacks responses or the channel control is from another population, withhold the calibrated estimate and explain why. Calibration is a model-assisted adjustment, not validation by itself.
Implementation
def queue_weighted_clarity(queue_outcomes, frame_counts):
weighted_yes = 0.0
represented = 0.0
for queue, (top_two, answered) in queue_outcomes.items():
if queue not in frame_counts or not 0 <= top_two <= answered or answered == 0:
raise ValueError("invalid queue outcomes")
respondent_weight = frame_counts[queue] / answered
weighted_yes += top_two * respondent_weight
represented += answered * respondent_weight
if set(queue_outcomes) != set(frame_counts):
raise ValueError("missing queue")
return weighted_yes / represented
share = queue_weighted_clarity({"standard": (28, 48), "partner": (12, 32)},
{"standard": 800, "partner": 400})
assert round(share, 3) == 0.514Performance and operating cost
Queue aggregation costs O(G) for G groups. Record-level weighting and audit checks cost O(R + I) over R responses and I invitations. Calibration adds iterative passes. The key unresolved risk is outcome-dependent response within adjustment cells.
Common Mistakes
- Do not use final weights without their frame and adjustment history.
- Do not report the weighted share as unbiased without the response assumption.
- Do not publish a calibrated subgroup rate with no respondent support.
