For prediction, fit replacement values on training data and preserve whether a value was missing at decision time.
Train-only imputation and missingness indicators
Decide which time the feature belongs to
A parcel-delay model is scored when a parcel leaves the depot. Its features may include a promised transit time and the carrier’s last known operating load. A final-delivery scan is a future outcome and cannot become an input, filled or not. Define the prediction timestamp first. Leakage checks apply to preprocessing as well as to the model itself.
Fit on training observations only
Suppose known depot queue times in training are 31, 47 and 55 minutes, while one training row is blank. Their median is 47. Store that value from the training fold, apply it to its blank row, and use the same stored value for validation and later production rows. If the validation batch changes the median, the model has learned from data reserved for evaluation. In cross-validation, each fold needs its own fitted imputer.
Keep an absence flag
Replacing a blank queue time with 47 makes it numerically identical to a measured 47. A separate flag tells the model which value was substituted. That flag may help if an offline scanner predicts service delays, but it can also encode a temporary incident that disappears after repair. Inspect missing rates by time and group before deploying the feature; the flag is a measured signal, not a causal explanation.
Use a baseline and compare
Median replacement is a simple baseline, not a universal distribution model. It compresses variation and can distort relationships when many values are absent. Compare a model with and without the missingness flag, and compare both against a policy that removes the affected feature. Evaluate on a time-held-out set representing future operations. Row deletion changes the evaluated population and may make scores look better for the wrong reason.
Store the fitted policy
Version the training cutoff, eligible feature columns, imputation statistic and absent-value definition. Apply exactly that transformation at scoring time. If a carrier begins sending an explicit string such as "unknown" instead of a null, normalize it before the stored transformation and measure the rate change. A pipeline prevents accidental refitting but cannot repair a badly defined prediction timestamp.
Implementation
from statistics import median
def fit_queue_time_policy(training_minutes):
measured = [minutes for minutes in training_minutes if minutes is not None]
if not measured:
raise ValueError("no measured training values")
return median(measured)
def transform_queue_time(values, replacement):
return [(replacement if value is None else value, value is None)
for value in values]
replacement = fit_queue_time_policy([31, 47, None, 55])
assert replacement == 47
assert transform_queue_time([None, 47, 61], replacement) == [
(47, True), (47, False), (61, False)]Performance and operating cost
Computing a median by sorting T measured training values costs O(T log T) time and O(T) memory. Transforming N rows costs O(N) time and output space. Persist the fitted statistic; recalculating it on a test or live batch changes the model contract.
Common Mistakes
- Do not calculate replacement values from validation or test data.
- Do not use an outcome that is unavailable at prediction time as a feature.
- Do not assume an absence flag remains predictive after collection systems change.
