Before filling a blank field, find out whether the observation process depends on the event or value being measured.
Missingness mechanisms: model why a value is absent
Define the record and the gap
A parcel-delivery table expects a final scan for every dispatched parcel. For 47 parcels, nine lack that scan. The missing value may mean an offline scanner, an unintegrated carrier, a parcel still in transit, or a genuine failure to scan. These are different states. Preserve the raw event stream and define a cutoff time before marking a scan absent. A field that is not yet due is not missing in the same way as a lost event.
Distinguish mechanisms carefully
Missing completely at random means scan absence is independent of observed and unobserved delivery facts. Missing at random means observed facts, such as carrier and depot, can explain the absence after conditioning on them. Missing not at random means the chance of a missing scan still depends on an unobserved outcome, such as whether the parcel was actually late. These are assumptions about a data-generating process, not labels that a null-count query can prove.
Measure patterns at the right grain
Count eligible records, missing records and missing rates by carrier, depot, calendar day and device version. Keep the denominator in every group. A jump from one missing scan among five parcels to three among six is more operationally meaningful than an unqualified count increase. Coverage errors can remove parcels before they reach this table, so inspect upstream counts too.
Separate causes from correlates
A high missing rate at one depot is evidence for investigation, not proof that the depot caused it. The depot may route a particular carrier whose scan events arrive late. Build a timeline of expected event, ingestion time and analysis cutoff, then ask the system owner about failures and retries. Do not infer an MCAR, MAR or MNAR mechanism from a significance test alone; several mechanisms can fit the same observed data.
Write a policy before analysis
Record which blank states are recoverable, which rows remain eligible for the decision, and which assumptions require sensitivity analysis. For a service report, a late scan can be treated as pending until a fixed grace window closes. Keep original missingness flags even if a later workflow fills values. Complete-case selection changes who is represented in the estimate.
Implementation
from collections import defaultdict
def scan_gap_rates(parcel_rows):
counts = defaultdict(lambda: {"eligible": 0, "missing": 0})
for parcel in parcel_rows:
if not parcel["scan_due"]:
continue
group = counts[parcel["carrier"]]
group["eligible"] += 1
group["missing"] += parcel["final_scan_at"] is None
return {carrier: (count["missing"], count["eligible"])
for carrier, count in counts.items()}
parcels = [
{"carrier": "North", "scan_due": True, "final_scan_at": None},
{"carrier": "North", "scan_due": True, "final_scan_at": "18:42"},
{"carrier": "South", "scan_due": False, "final_scan_at": None},
]
assert scan_gap_rates(parcels) == {"North": (1, 2)}Performance and operating cost
A single pass over N parcel rows costs O(N) time and O(G) space for G carrier groups. The result reports observed patterns; it does not establish a missingness mechanism. Event-history joins may dominate runtime when the source stream is large.
Common Mistakes
- Do not classify a not-yet-due scan as missing.
- Do not infer a mechanism solely from an observed missing-rate table.
- Do not overwrite the original absence flag when repairing data.
