A missing value records unavailable information, not a numeric zero or a failed outcome; the reason and timing of absence determine how to analyze it.
Missing data policy: distinguish absence from a measured zero
Classify the absence
For a receipt dataset, amount_paid can be zero because no payment occurred, null because a payment system had no record, or null because the payment event had not arrived by the analysis cutoff. These are different states. Count missing values by source, week and cohort before filling them. A blanket dropna may remove the hardest cases and bias the estimate. A blanket fillna(0) can turn ingestion failure into a false business fact.
Preserve a signal and an audit trail
Create an explicit missing indicator when absence is informative, but do not let it smuggle future information into a predictive feature. If the amount is needed for modeling, estimate an imputation value from the training partition only and apply that same learned value to validation and serving data. For reporting, it may be more honest to show an unknown bucket and a coverage percentage than to impute at all. The preprocessing pipeline] keeps training and evaluation boundaries intact.
Check sensitivity
Compute the target metric on observed records, then under a conservative lower and upper scenario for missing records. If the conclusion changes, the missingness policy is a decision risk, not a housekeeping detail. Document the cutoff, the number of affected rows and whether missingness differs across groups. The denominator contract] must say whether unknown outcomes count as pending, excluded or failed; never let pandas choose that policy by default.
Implementation
receipt_facts["amount_missing"] = receipt_facts["amount_paid"].isna()
coverage = receipt_facts.groupby("submission_week", dropna=False).agg(
submissions=("submission_id", "size"),
missing_amounts=("amount_missing", "sum"))
coverage["observed_share"] = (
1 - coverage["missing_amounts"] / coverage["submissions"])
assert coverage["observed_share"].between(0, 1).all()Performance and operating cost
A column scan is O(N) time and a boolean indicator needs O(N) extra space. Grouped diagnostics need memory proportional to the number of groups; imputation can be cheap computationally while still introducing large inferential error.
Common Mistakes
- Do not equate null with zero.
- Do not drop incomplete rows before comparing their cohort mix.
- Do not learn imputation values from the evaluation set.
Read next
- Dataset grain and join cardinality: protect the unit of analysis
- Metric denominators and cohorts: make a rate reproducible
- Exploratory analysis without peeking: inspect the data and preserve the test
- Leakage-safe preprocessing: fit every learned transform inside the training fold
Connected implementation
- Pandas nullable integers: missing amounts are not zero
- Pandas groupby(dropna=False): keep a missing category visible
Continue the workflow: Time-series calendar: distinguish missing periods from measured zeros.
Continue the workflow: Nonresponse bias: diagnose the missing outcomes before adjusting.
Continue the workflow: Right censoring and time-to-event analysis for open cases.
Continue the workflow: SQL null semantics: count records without inventing outcomes.
Continue the workflow: Missingness mechanisms: model why a value is absent.
Continue the workflow: Response options and ordinal coding without false precision.
