Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Survey calibration: distinguish known joint cells from separate margins

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Poststratification can match known joint population cells; raking iteratively matches separate known margins when their joint table is unavailable.

Start with the survey design

A contact-center survey oversamples high-volume branches and loses more respondents from late shifts. Begin with inclusion probabilities and a response ledger for the full invited frame. Design weights account for unequal selection; adjustment for nonresponse needs assumptions tied to recorded variables. Calibration then aligns weighted sample totals with independently known population counts, such as branch-by-shift cells or separate branch and shift margins. The survey-estimate lesson establishes what population the weights target.

Use joint cells when known

If every branch-by-shift population count is reliable and represented in the sample, multiply each base weight by target cell total divided by the cell’s current weighted total. The code performs that direct poststratification for one declared cell label. A target cell with positive population count but no sample support cannot be created by multiplying weights; it needs a broader grouping or an explicit model. A zero target with sampled records also deserves a frame audit before those records are silently assigned zero weight.

Rake when only margins are known

When the joint cross-tabulation is unavailable, iterative proportional fitting alternates adjustments for branch and shift totals until both sets of margins agree within tolerance. This does not ensure the joint distribution is correct, nor does it repair nonresponse driven by unmeasured dissatisfaction within cells. The margins should describe the same population and time window; inconsistent totals cannot all be matched. The weight-diagnostic lesson checks whether calibration concentrates influence.

Carry the design into uncertainty

Use variance estimation that reflects sampling clusters, stratification, and the calibration procedure. A simple row-independent standard error for a weighted average is usually incomplete. Report base and final weight ranges, target margins, unsupported cells, convergence tolerance if raking, and sensitivity to trimming. The project gates a branch satisfaction headline when calibration support or population totals are inconsistent.

Implementation

python
def poststratify_cell_weights(sample_rows, target_counts):
    current = {cell: 0.0 for cell in target_counts}
    for cell, base_weight in sample_rows:
        if cell not in target_counts or base_weight <= 0:
            raise ValueError("known cell and positive base weight required")
        current[cell] += base_weight
    if any(total > 0 and current[cell] == 0
           for cell, total in target_counts.items()):
        raise ValueError("positive target cell lacks sample support")
    return [(cell, base_weight * target_counts[cell] / current[cell])
            for cell, base_weight in sample_rows]

weights = poststratify_cell_weights([("day", 2), ("day", 2), ("late", 1)],
                                   {"day": 30, "late": 18})
assert [weight for _, weight in weights] == [15, 15, 18]

Performance and operating cost

Direct poststratification is O(n + g) time and O(g) space for n sample records and g cells. Raking adds repeated passes over variables until convergence. Neither procedure manufactures observations for an unsupported group or removes bias from unmeasured nonresponse.

Common Mistakes

  • Calibrating to margins from a different population or date.
  • Assigning weights to an empty positive target cell as if observations existed.
  • Calling separate-margin raking equivalent to a known joint distribution.
  • Using final weights with a row-independent variance formula that ignores the design.

Read next

ai-data
applied-statistics
Storage details