Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Data science begins with a decision, population and clock

Last updated: 5 Oct 20265 min read
concept
BeginnerBy AITrove Editorial

A useful analysis starts by naming the decision it supports, the population it covers and the exact time at which information is available.

Define the question as a contract

Suppose a receipt team wants to reduce delayed reviews. “Analyze receipts” is too broad. Ask which action might change: staffing, routing or a customer message. Specify the target population, the submission window, the outcome definition and who will act on the result. Decide what a false claim would cost. This prevents a visually interesting correlation from becoming a policy without evidence.

Inventory the data before writing a model

List each source, row grain, stable key, event timestamp, ingestion delay, owner and access rule. A submission table and a review-events table answer different questions; a careless join changes the number of rows. Grain and cardinality] come before charts. Check which fields are available at the moment a live decision would be made, especially if this analysis may become a prediction service.

Establish a simple baseline

Count eligible submissions and delayed reviews by week. Inspect missing timestamps, duplicate keys and extreme delays. Report numerator and denominator together. If a weekly rate moves, distinguish an actual operational change from a partial week or late arriving outcomes. Cohort definitions] make that distinction testable.

Deliver a reproducible result

Record the snapshot, cutoff, transformation revision and known limitations. Show a decision-maker the baseline and a concrete uncertainty: “47 of 311 eligible submissions exceeded 24 hours; 12 lack a final review timestamp.” That sentence is more useful than a dashboard with an unexplained percentage. Then decide whether descriptive analysis is enough or a model would change the action.

Implementation

python
analysis_contract = {
    "decision": "assign next-day review capacity",
    "unit": "one submitted receipt",
    "window": "submission timestamp in UTC",
    "outcome": "review completed within 24 hours",
    "cutoff": "2026-09-27T12:00:00Z",
}
assert analysis_contract["unit"] == "one submitted receipt"

Performance and operating cost

A simple baseline is usually O(N) over the defined population. The costly work is obtaining a trustworthy snapshot and agreeing on the outcome before optimizing code.

Common Mistakes

  • Do not start with an algorithm before naming the decision.
  • Do not mix event rows with submission rows in a rate.
  • Do not conceal incomplete outcome data behind one percentage.

Read next

Related path: Prediction-time feature availability: reject future information before training.

Continue with: Data science core concepts: grain, distributions, uncertainty and action.

Continue the workflow: Sampling frames and coverage error: who could enter the analysis?.

data-science
Storage details