A risk set contains only units still eligible just before an event time; a landmark prediction resets the question for units that remain event-free at that landmark.
Survival risk sets and landmark snapshots
Keep the event order explicit
At day 28 after inspection, a vehicle that already had a replacement is no longer at risk for its first replacement. A vehicle censored at day 19 is also absent because follow-up ended. A vehicle censored at day 43 was known to remain event-free through day 28 and belongs in the risk set. This distinction drives survival likelihoods and evaluation denominators.
Treat delayed entry separately
A vehicle may join the fleet monitoring feed 12 days after inspection. If it was eligible but invisible before entry, placing it in earlier risk sets grants unobserved event-free time. Record entry age and include the vehicle only after entry. Delayed entry is not the same as right-censoring at the end. Entry bias can distort comparisons.
Build landmark datasets for fresh decisions
At day 30, operations may ask for risk over the next 45 days using the current voltage reading. Include only vehicles still event-free and observed at day 30; use readings available by day 30. The target begins at the landmark, not the original inspection. A new landmark creates a new prediction episode and requires grouped splitting by vehicle.
State endpoint conventions
If an event is recorded at exactly day 30, decide whether it belongs before or after a day-30 snapshot. Use one consistent half-open interval convention in data creation and model evaluation. The code uses entry <= t and exit >= t for a just-before-t risk set; events at t are included until that event is processed.
Check support by slice
A far-future landmark may retain only well-maintained vehicles, changing the population. Report risk-set size by depot and vehicle model at each candidate horizon. Horizon evaluation should not hide a tiny late risk set.
Implementation
follow_up = [
{"vehicle": "fleet-47", "entry": 0, "exit": 57, "event": True},
{"vehicle": "fleet-62", "entry": 12, "exit": 71, "event": False},
{"vehicle": "fleet-83", "entry": 0, "exit": 28, "event": True},
{"vehicle": "fleet-94", "entry": 0, "exit": 19, "event": False},
]
def at_risk_just_before(records, day):
return [row["vehicle"] for row in records
if row["entry"] <= day <= row["exit"]]
assert at_risk_just_before(follow_up, 28) == [
"fleet-47", "fleet-62", "fleet-83"]
assert at_risk_just_before(follow_up, 58) == ["fleet-62"]Performance and operating cost
A direct scan costs O(N) time per event or landmark and O(N) output memory. Sorting entry and exit times supports faster repeated queries; full model fitting may need event-time risk summaries. Landmark tables increase row count and require vehicle-level grouping to keep evaluation independent.
Common Mistakes
- Do not place a vehicle into risk sets before it entered observation.
- Do not let a prior event remain at risk for its first occurrence.
- Do not treat a landmark dataset as independent rows during splitting.
