Data science begins with a decision, a population and an observation clock. Learn the shape and quality of the data before making a claim.
Foundations
- Data science begins with a decision, population and clock
- Data science core concepts: grain, distributions, uncertainty and action
Data contracts, analysis and uncertainty
- Dataset grain and join cardinality: protect the unit of analysis
- Missing data policy: distinguish absence from a measured zero
- Exploratory analysis without peeking: inspect the data and preserve the test
- Metric denominators and cohorts: make a rate reproducible
- Reproducible analysis snapshots: pin data, code and cutoff together
- Bootstrap intervals: estimate uncertainty at the right sampling unit
Apply the work
Continue the workflow: Privacy units and bounded contribution: count people, not events.
Continue into another subject: Geospatial Analytics Tutorial.
Sampling and representativeness
- Sampling frames and coverage error: who could enter the analysis?
- Stratified sampling and design weights for operational estimates
- Nonresponse bias: diagnose the missing outcomes before adjusting
- Precision budgets and effective sample size for weighted analyses
Profiling and resilient analysis
- Column profiles and domain constraints before analysis
- Outlier investigation: impossible values, rare events and quarantines
- Resistant summaries: median, trimmed mean and median absolute deviation
- Subgroup distributions and aggregation reversal
Additional projects
- Project: estimate support-case defects from an auditable sample
- Project: audit fulfillment delays without hiding the tail
Observational evidence and decision sensitivity
- Cohort entry and survivorship bias: count the cases that could fail
- Right censoring and time-to-event analysis for open cases
- Proxy measurements and label error in operational datasets
- Sensitivity analysis: find which assumptions can reverse a decision
Support resolution project
SQL analysis and metric verification
- SQL null semantics: count records without inventing outcomes
- SQL conditional aggregation: rates with an auditable denominator
- SQL window functions: select one event with a deterministic rule
- SQL reconciliation: prove the cohort survived each transformation
SQL audit project
Product behavior and retention analysis
- Product event contracts: identity, deduplication and eligibility
- Ordered funnels: step order, entry rules and conversion windows
- Retention cohorts: return behavior only after a full observation window
- Instrumentation changes and metric guardrails for product analysis
Checkout audit project
Measurement quality and uncertainty
- Quantity units and conversion contracts for mixed datasets
- Measurement resolution and rounding: avoid invented precision
- Calibration drift: compare sensor readings with a reference
- Derived metrics: carry input bounds and shared error into the result
Cold-chain audit project
Record linkage and entity resolution
- Identifier normalization and exact linkage without false merges
- Blocking for record linkage: reduce comparisons without hiding true matches
- Match scores and review bands: separate similarity from identity
- Entity merge lineage: show how identity decisions change metrics
Customer identity project
Decision optimization and allocation
- Decision objectives: define the action before optimizing a score
- Constraint feasibility and slack: reject impossible plans early
- Integer allocation and relaxation gaps for indivisible work
- Scenario stress tests and regret for uncertain allocation benefits
Support allocation project
Missing observations and valid inference
- Missingness mechanisms: model why a value is absent
- Complete-case selection: know whose outcome remains
- Train-only imputation and missingness indicators
- Missing-not-at-random sensitivity: test what unseen outcomes could change
Delivery scan audit project
Bayesian event-rate analysis
- Prior predictive checks for a binary service rate
- Beta-binomial updating with an auditable case count
- Posterior intervals and threshold decisions
- Shared priors, shrinkage and the limits of fixed pooling
Bayesian escalation project
Survey measurement and form quality
- Survey constructs, response units and recall windows
- Response options and ordinal coding without false precision
- Cognitive pilots and survey branch-logic audits
- Question versions, order effects and comparable trends
Support handoff survey project
Bayesian model checks and monitoring
- Posterior predictive checks for a hidden group gap
- Prior sensitivity for an operational action threshold
- Held-out predictive log scores for a binary rate model
- Sequential posterior monitoring with mature outcomes
Support rate model-checking project
Survey weighting and population estimates
- Survey inclusion probabilities and base weights
- Nonresponse adjustment cells and the positivity check
- Survey calibration to known population margins
- Weight extremes, effective sample size and trimming tradeoffs
Support survey weighting project
Survival analysis and competing outcomes
- Event-time risk sets and tied outcomes
- Kaplan–Meier curves for support resolution
- Competing risks and resolution incidence
- Restricted mean unresolved time at a business horizon
Support resolution-time project
Advanced time-to-event decisions
- Delayed entry and left-truncated risk sets
- Survival-curve uncertainty and thin-tail support
- Landmark analysis and immortal-time traps
- Compare queues by restricted mean open time
Specialist-queue survival project
Monte Carlo decisions and simulation quality
- Monte Carlo estimands and reproducible draws
- Simulation error and replication budgets
- Joint inputs and correlated operational shocks
- Paired policy simulations with common random numbers
Warehouse capacity simulation project
Process monitoring and alert quality
- Control limits and service targets answer different questions
- P-charts with changing daily denominators
- EWMA monitoring for small persistent rate shifts
- Frozen baselines and alert investigation
Support breach monitoring project
Survival models and prediction checks
- Cox partial likelihood and hazard-ratio interpretation
- Diagnose the proportional-hazards assumption
- Interval-censored support resolution times
- Calibrate resolution predictions at a fixed horizon
Support resolution model review
Longitudinal histories and survival validation
- Time-updated case histories and the observation clock
- Counting-process intervals for a time-varying Cox model
- Censoring survival weights and horizon support
- IPCW Brier score for censored survival predictions
- Temporal validation of a support-resolution model
Longitudinal resolution validation project
Competing exits and prediction uncertainty
- Competing-outcome contracts at a fixed horizon
- Aalen–Johansen cumulative incidence from case events
- From cause-specific event rates to absolute risk
- Calibrate competing-risk predictions at a fixed horizon
- Paired bootstrap uncertainty for prediction-score differences
Competing exit prediction review
Staggered rollout analysis
- Branch-month panels and the rollout clock
- Cohort-time difference-in-differences for staggered adoption
- Event-time pre-trends and support for a staggered rollout
- Spillover-aware comparison pools for branch rollouts
- Branch-cluster bootstrap for a rollout contrast
Staggered branch rollout audit
Small rollout inference and spillovers
- Assignment-aware permutation tests for a few branches
- Leave-one-branch-out stability for rollout effects
- Direct and indirect exposure maps for branch rollouts
- A spillover difference-in-differences contrast
- Negative-control outcomes and placebo rollout checks
Small rollout and spillover review
Synthetic controls and effect transfer
- Synthetic-control donor eligibility and pre-fit
- Synthetic-control weight search and prediction
- Synthetic-control placebos and fit ratios
- Synthetic-control donor and window sensitivity
- Target-population standardization for branch effects
Single-branch policy transfer review
Interrupted time-series policy evaluation
- Interrupted-series outcome clock and data contract
- Segmented regression for level and slope changes
- Controlled interrupted time series with a comparison series
- Seasonality and residual dependence in interrupted series
- Transition windows and lag sensitivity for policy series
Interrupted-series policy review
Outcome-label error and validation
- Design a validation subsample for outcome labels
- Estimate outcome-label sensitivity and specificity by group
- Correct a binary outcome rate for label error
- Differential outcome-label error in group contrasts
- Misclassification assumption grids and decision ranges
Outcome-label error audit
Partial identification and outcome bounds
- Partial identification and an assumption ledger
- Worst-case bounds for missing binary outcomes
- Treatment-effect bounds under differential attrition
- Bounded missing-outcome risk scenarios
- Decision thresholds and follow-up value under bounds
