Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Missing data policy: distinguish absence from a measured zero

Last updated: 7 Oct 20265 min read
tutorial
BeginnerBy AITrove Editorial

A missing value records unavailable information, not a numeric zero or a failed outcome; the reason and timing of absence determine how to analyze it.

Classify the absence

For a receipt dataset, amount_paid can be zero because no payment occurred, null because a payment system had no record, or null because the payment event had not arrived by the analysis cutoff. These are different states. Count missing values by source, week and cohort before filling them. A blanket dropna may remove the hardest cases and bias the estimate. A blanket fillna(0) can turn ingestion failure into a false business fact.

Preserve a signal and an audit trail

Create an explicit missing indicator when absence is informative, but do not let it smuggle future information into a predictive feature. If the amount is needed for modeling, estimate an imputation value from the training partition only and apply that same learned value to validation and serving data. For reporting, it may be more honest to show an unknown bucket and a coverage percentage than to impute at all. The preprocessing pipeline] keeps training and evaluation boundaries intact.

Check sensitivity

Compute the target metric on observed records, then under a conservative lower and upper scenario for missing records. If the conclusion changes, the missingness policy is a decision risk, not a housekeeping detail. Document the cutoff, the number of affected rows and whether missingness differs across groups. The denominator contract] must say whether unknown outcomes count as pending, excluded or failed; never let pandas choose that policy by default.

Implementation

python
receipt_facts["amount_missing"] = receipt_facts["amount_paid"].isna()
coverage = receipt_facts.groupby("submission_week", dropna=False).agg(
    submissions=("submission_id", "size"),
    missing_amounts=("amount_missing", "sum"))
coverage["observed_share"] = (
    1 - coverage["missing_amounts"] / coverage["submissions"])
assert coverage["observed_share"].between(0, 1).all()

Performance and operating cost

A column scan is O(N) time and a boolean indicator needs O(N) extra space. Grouped diagnostics need memory proportional to the number of groups; imputation can be cheap computationally while still introducing large inferential error.

Common Mistakes

  • Do not equate null with zero.
  • Do not drop incomplete rows before comparing their cohort mix.
  • Do not learn imputation values from the evaluation set.

Read next

Connected implementation

Continue the workflow: Time-series calendar: distinguish missing periods from measured zeros.

Continue the workflow: Nonresponse bias: diagnose the missing outcomes before adjusting.

Continue the workflow: Right censoring and time-to-event analysis for open cases.

Continue the workflow: SQL null semantics: count records without inventing outcomes.

Continue the workflow: Missingness mechanisms: model why a value is absent.

Continue the workflow: Response options and ordinal coding without false precision.

data-science
missing-data-policy
Storage details