A useful analysis starts by naming the decision it supports, the population it covers and the exact time at which information is available.
Data science begins with a decision, population and clock
Define the question as a contract
Suppose a receipt team wants to reduce delayed reviews. “Analyze receipts” is too broad. Ask which action might change: staffing, routing or a customer message. Specify the target population, the submission window, the outcome definition and who will act on the result. Decide what a false claim would cost. This prevents a visually interesting correlation from becoming a policy without evidence.
Inventory the data before writing a model
List each source, row grain, stable key, event timestamp, ingestion delay, owner and access rule. A submission table and a review-events table answer different questions; a careless join changes the number of rows. Grain and cardinality] come before charts. Check which fields are available at the moment a live decision would be made, especially if this analysis may become a prediction service.
Establish a simple baseline
Count eligible submissions and delayed reviews by week. Inspect missing timestamps, duplicate keys and extreme delays. Report numerator and denominator together. If a weekly rate moves, distinguish an actual operational change from a partial week or late arriving outcomes. Cohort definitions] make that distinction testable.
Deliver a reproducible result
Record the snapshot, cutoff, transformation revision and known limitations. Show a decision-maker the baseline and a concrete uncertainty: “47 of 311 eligible submissions exceeded 24 hours; 12 lack a final review timestamp.” That sentence is more useful than a dashboard with an unexplained percentage. Then decide whether descriptive analysis is enough or a model would change the action.
Implementation
analysis_contract = {
"decision": "assign next-day review capacity",
"unit": "one submitted receipt",
"window": "submission timestamp in UTC",
"outcome": "review completed within 24 hours",
"cutoff": "2026-09-27T12:00:00Z",
}
assert analysis_contract["unit"] == "one submitted receipt"Performance and operating cost
A simple baseline is usually O(N) over the defined population. The costly work is obtaining a trustworthy snapshot and agreeing on the outcome before optimizing code.
Common Mistakes
- Do not start with an algorithm before naming the decision.
- Do not mix event rows with submission rows in a rate.
- Do not conceal incomplete outcome data behind one percentage.
Read next
- Dataset grain and join cardinality: protect the unit of analysis
- Metric denominators and cohorts: make a rate reproducible
- Reproducible analysis snapshots: pin data, code and cutoff together
- Machine learning starts with a baseline and a valid prediction clock
Related path: Prediction-time feature availability: reject future information before training.
Continue with: Data science core concepts: grain, distributions, uncertainty and action.
Continue the workflow: Sampling frames and coverage error: who could enter the analysis?.
