Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Sampling frames and coverage error: who could enter the analysis?

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A sampling frame is the operational list from which observations can be selected; it may not cover the population the decision concerns.

Specify the population first

A support team wants the defect rate among all submitted cases last quarter. Its export contains only cases assigned to agents, omitting self-service closures and bounced submissions. Even a random draw from that export cannot recover the absent units. Define the case identity, eligibility window, decision time and outcome before querying. Dataset grain prevents one case with several status events from counting several times.

Audit the frame

Compare eligible intake IDs with the export. Count missing IDs, duplicate IDs, unknown eligibility and records that arrived after the cutoff. Break the differences down by channel and week. Coverage error is not the same as random sampling error: a larger draw from a systematically incomplete list can make the estimate more precise around the wrong quantity. The estimand defines what the final number is allowed to describe.

Record selection before outcomes

Freeze the frame version and selection mechanism before viewing labels. If high-severity cases have a separate queue, include that queue or limit the claim to the ordinary queue. Capture inclusion probability for each selected unit and a reason for excluded units. Later joins to reviewer decisions should use stable case IDs and an as-of cutoff, not a mutable current-status table.

Rehearse a discrepancy

Suppose 1,500 intake IDs appear in the event log, but only 1,320 are in the agent export. Of the 180 missing, 145 are self-service closures and 35 failed assignment. A defect estimate from the export applies to assigned cases unless those groups are separately measured or brought into the frame. Report the 180-case gap before quoting a rate.

Implementation

python
def coverage_gaps(eligible_case_ids, frame_case_ids):
    eligible = set(eligible_case_ids)
    frame = set(frame_case_ids)
    return {"missing": eligible - frame,
            "out_of_scope": frame - eligible,
            "covered": eligible & frame}

Performance and operating cost

Building two ID sets costs O(N + M) expected time and space for N eligible IDs and M frame IDs. The result checks coverage by identity; deciding eligibility and resolving duplicate business entities require separate rules.

Common Mistakes

  • Do not call a random sample representative when the frame excludes a channel.
  • Do not use current mutable status to reconstruct last quarter’s eligibility.
  • Do not count status events as distinct cases.

Read next

Continue the workflow: Complete-case selection: know whose outcome remains.

ai-data
data-science
Storage details