Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Data science core concepts: grain, distributions, uncertainty and action

Last updated: 5 Oct 20265 min read
concept
BeginnerBy AITrove Editorial

Data science connects a defined population and trustworthy measurements to a decision while accounting for selection, missingness and uncertainty.

Observation is not the target population

A dataset contains observed rows, not necessarily every case the decision will affect. A receipt report built only from successfully processed submissions excludes the failures operations needs to understand. Name the eligible population and selection path. Quantify what was dropped before interpreting a mean, chart or model score.

Summary depends on shape

A mean delay can be pulled upward by a few extreme cases. Report the median and a high quantile when tail latency matters, with counts and the observation window. Compare groups only when they share a compatible definition and enough independent units. Exploration] should produce checks and questions, not a hunt for a pleasing graph.

Association is not intervention

A busy queue and slower reviews may move together because load affects both. Moving receipts to another queue does not necessarily solve the underlying staffing constraint. Before suggesting a causal intervention, identify confounders, timing and a credible comparison design. For prediction, ask a different question: will the features be available when the decision is made? Availability audits] address that boundary.

Show uncertainty honestly

Sampling error is one uncertainty, but delayed data, missing values and policy changes often dominate. Write these separately. Bootstrap intervals] estimate variation only under the resampling design; they do not correct bad source data. A useful result ends with a recommended decision, the evidence supporting it and the condition that would overturn it.

Implementation

python
summary = {
    "eligible": len(eligible_receipts),
    "observed_delay": int(eligible_receipts["delay_hours"].notna().sum()),
    "median_delay": eligible_receipts["delay_hours"].median(),
    "p90_delay": eligible_receipts["delay_hours"].quantile(0.9),
}
assert summary["observed_delay"] <= summary["eligible"]

Performance and operating cost

Grouped scans and exact quantiles can consume substantial memory on large data. A small sample is useful for inspection, but final population claims require the defined full extract or an explicit sampling design.

Common Mistakes

  • Do not confuse the observed table with the full eligible population.
  • Do not infer causation from correlation.
  • Do not present a narrow interval as proof the measurements are unbiased.

Read next

Study next: Dataset grain and join cardinality: protect the unit of analysis.

Study next: Missing data policy: distinguish absence from a measured zero.

Study next: Exploratory analysis without peeking: inspect the data and preserve the test.

Study next: Metric denominators and cohorts: make a rate reproducible.

Study next: Reproducible analysis snapshots: pin data, code and cutoff together.

Study next: Bootstrap intervals: estimate uncertainty at the right sampling unit.

data-science
Storage details