An analysis needs a target population, an exact quantity and a sampling process that can observe that population.
Population, estimand and sampling frame: name the quantity before calculating
State the target
A receipt-review rate might mean all receipts submitted in September, only those eligible for manual review, or only those with a completed decision. These are different populations. Define the event window, eligibility rule, unit and numerator before writing a query. Metric denominators] turn a vague business question into a reproducible count.
Describe the observation path
The sampling frame is the set of units that could enter the dataset. If mobile submissions are missing during an outage, the observed frame omits them; a tiny standard error cannot repair that bias. Record coverage by channel, region and time. A convenience sample of reviewed receipts can estimate properties of reviewed receipts, but not necessarily all submissions.
Check dependence
One customer may submit several receipts, and one receipt may create multiple events. Decide whether the unit is customer, receipt or event. If units are clustered, treating every row as independent understates uncertainty. Clustered uncertainty] keeps repeated observations together.
Make exclusions visible
Produce a flow of counts from raw submissions to eligible units, observed outcomes and final analysis rows. Reconcile missing IDs and late outcomes. Test a duplicated receipt and an outage day. If the sample changes because a late outcome arrives, version the analysis cutoff rather than silently replacing yesterday’s denominator.
Implementation
def eligible_receipts(records, window_start, window_end):
selected = {}
for receipt in records:
if not window_start <= receipt.submitted_at < window_end:
continue
if receipt.receipt_id in selected:
raise ValueError("receipt appears more than once")
selected[receipt.receipt_id] = receipt
return list(selected.values())Performance and operating cost
A scan is O(N) over source rows; a dictionary for uniqueness uses O(K) memory for K eligible receipts. Recording exclusions adds storage but prevents a precise estimate of the wrong population.
Common Mistakes
- Do not treat a convenience sample as the entire target population.
- Do not count receipt events as independent receipts.
- Do not let late outcomes change a historical denominator without a version.
Read next
- Distribution summaries: report tails and define the outlier policy
- Standard error and cluster bootstrap: resample the independent unit
- Metric denominators and cohorts: make a rate reproducible
- Dataset grain and join cardinality: protect the unit of analysis
Continue the workflow: Causal estimands: name the unit, treatment and missing counterfactual.
Continue the workflow: Sampling frames and coverage error: who could enter the analysis?.
Continue the workflow: Linear regression slopes: define the comparison a coefficient makes.
Continue the workflow: Stratified survey estimates: weight toward the named population.
Continue the workflow: Event counts and exposure: compare rates across unequal observation time.
Continue the workflow: Finite-frame sampling: adjust uncertainty for sampling without replacement.
