Changing covariates belong to intervals where their values were already known; copying the latest value backward creates a future-information leak.
Time-varying survival features without future leakage
Choose a clock for each measurement
Fleet voltage is measured after inspections and maintenance visits. A reading taken on day 41 can inform predictions made after day 41, not the day-24 inspection risk. Record observation and availability timestamps, since uploads may arrive after measurements. Use the time when the service could have used the value.
Represent stable intervals
A time-varying model can split a vehicle episode into start-stop intervals. Each interval carries the value known before its endpoint; status is true only on the interval ending with the event. Intervals cannot overlap or extend beyond the observed exit. Risk sets then use the current value at each event time.
Separate interventions from predictors
A new battery-warning flag may cause maintenance that prevents replacement. Treating that flag as an ordinary predictor can entangle prediction with an intervention policy. If the aim is counterfactual risk under no action, a standard observational survival model is not enough. State the actual estimand: risk under the current operational policy.
Avoid backfill artifacts
A warehouse may overwrite a voltage field with its latest value. Historical snapshots must come from an append-only log or versioned table. Forward filling is acceptable only from earlier available readings and with a maximum age rule; never fill backward from a later inspection. Availability auditing applies at interval level.
Test the serving snapshot
Replay a past inspection and a day-45 landmark from timestamped records. Confirm that every emitted feature existed at that decision time. Compare a baseline-only model to a landmark update model on future vehicles, using paired horizons and the same censoring policy.
Implementation
readings = [
{"available_day": 24, "voltage": 12.6},
{"available_day": 41, "voltage": 12.2},
{"available_day": 58, "voltage": 11.9},
]
def last_available_voltage(observations, decision_day):
eligible = [reading for reading in observations
if reading["available_day"] <= decision_day]
return max(eligible, key=lambda reading: reading["available_day"])["voltage"] if eligible else None
assert last_available_voltage(readings, 30) == 12.6
assert last_available_voltage(readings, 45) == 12.2
assert last_available_voltage(readings, 20) is NonePerformance and operating cost
The simple snapshot scan is O(M) per request over M readings; sorted, indexed histories support O(log M) lookup. Interval expansion adds storage proportional to measurement changes. Event-time fitting may be substantially more expensive than a baseline-only model, so measure whether updated predictions improve decisions.
Common Mistakes
- Do not copy the final voltage reading into earlier intervals.
- Do not confuse measurement timestamp with availability timestamp.
- Do not interpret policy-affected risk as untreated failure risk.
