Data science connects a defined population and trustworthy measurements to a decision while accounting for selection, missingness and uncertainty.
Data science core concepts: grain, distributions, uncertainty and action
Observation is not the target population
A dataset contains observed rows, not necessarily every case the decision will affect. A receipt report built only from successfully processed submissions excludes the failures operations needs to understand. Name the eligible population and selection path. Quantify what was dropped before interpreting a mean, chart or model score.
Summary depends on shape
A mean delay can be pulled upward by a few extreme cases. Report the median and a high quantile when tail latency matters, with counts and the observation window. Compare groups only when they share a compatible definition and enough independent units. Exploration] should produce checks and questions, not a hunt for a pleasing graph.
Association is not intervention
A busy queue and slower reviews may move together because load affects both. Moving receipts to another queue does not necessarily solve the underlying staffing constraint. Before suggesting a causal intervention, identify confounders, timing and a credible comparison design. For prediction, ask a different question: will the features be available when the decision is made? Availability audits] address that boundary.
Show uncertainty honestly
Sampling error is one uncertainty, but delayed data, missing values and policy changes often dominate. Write these separately. Bootstrap intervals] estimate variation only under the resampling design; they do not correct bad source data. A useful result ends with a recommended decision, the evidence supporting it and the condition that would overturn it.
Implementation
summary = {
"eligible": len(eligible_receipts),
"observed_delay": int(eligible_receipts["delay_hours"].notna().sum()),
"median_delay": eligible_receipts["delay_hours"].median(),
"p90_delay": eligible_receipts["delay_hours"].quantile(0.9),
}
assert summary["observed_delay"] <= summary["eligible"]Performance and operating cost
Grouped scans and exact quantiles can consume substantial memory on large data. A small sample is useful for inspection, but final population claims require the defined full extract or an explicit sampling design.
Common Mistakes
- Do not confuse the observed table with the full eligible population.
- Do not infer causation from correlation.
- Do not present a narrow interval as proof the measurements are unbiased.
Read next
- Data science begins with a decision, population and clock
- Missing data policy: distinguish absence from a measured zero
- Exploratory analysis without peeking: inspect the data and preserve the test
- Bootstrap intervals: estimate uncertainty at the right sampling unit
Study next: Dataset grain and join cardinality: protect the unit of analysis.
Study next: Missing data policy: distinguish absence from a measured zero.
Study next: Exploratory analysis without peeking: inspect the data and preserve the test.
Study next: Metric denominators and cohorts: make a rate reproducible.
Study next: Reproducible analysis snapshots: pin data, code and cutoff together.
Study next: Bootstrap intervals: estimate uncertainty at the right sampling unit.
