A sampling frame is the operational list from which observations can be selected; it may not cover the population the decision concerns.
Sampling frames and coverage error: who could enter the analysis?
Specify the population first
A support team wants the defect rate among all submitted cases last quarter. Its export contains only cases assigned to agents, omitting self-service closures and bounced submissions. Even a random draw from that export cannot recover the absent units. Define the case identity, eligibility window, decision time and outcome before querying. Dataset grain prevents one case with several status events from counting several times.
Audit the frame
Compare eligible intake IDs with the export. Count missing IDs, duplicate IDs, unknown eligibility and records that arrived after the cutoff. Break the differences down by channel and week. Coverage error is not the same as random sampling error: a larger draw from a systematically incomplete list can make the estimate more precise around the wrong quantity. The estimand defines what the final number is allowed to describe.
Record selection before outcomes
Freeze the frame version and selection mechanism before viewing labels. If high-severity cases have a separate queue, include that queue or limit the claim to the ordinary queue. Capture inclusion probability for each selected unit and a reason for excluded units. Later joins to reviewer decisions should use stable case IDs and an as-of cutoff, not a mutable current-status table.
Rehearse a discrepancy
Suppose 1,500 intake IDs appear in the event log, but only 1,320 are in the agent export. Of the 180 missing, 145 are self-service closures and 35 failed assignment. A defect estimate from the export applies to assigned cases unless those groups are separately measured or brought into the frame. Report the 180-case gap before quoting a rate.
Implementation
def coverage_gaps(eligible_case_ids, frame_case_ids):
eligible = set(eligible_case_ids)
frame = set(frame_case_ids)
return {"missing": eligible - frame,
"out_of_scope": frame - eligible,
"covered": eligible & frame}Performance and operating cost
Building two ID sets costs O(N + M) expected time and space for N eligible IDs and M frame IDs. The result checks coverage by identity; deciding eligibility and resolving duplicate business entities require separate rules.
Common Mistakes
- Do not call a random sample representative when the frame excludes a channel.
- Do not use current mutable status to reconstruct last quarter’s eligibility.
- Do not count status events as distinct cases.
Read next
- Stratified sampling and design weights for operational estimates
- Nonresponse bias: diagnose the missing outcomes before adjusting
- Dataset grain and join cardinality: protect the unit of analysis
- Population, estimand and sampling frame: name the quantity before calculating
Continue the workflow: Complete-case selection: know whose outcome remains.
