Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Time-to-event targets and censoring contracts

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A survival target pairs elapsed time with event status; an unfinished observation contributes follow-up rather than an invented event or a negative outcome.

Choose one time origin

A fleet operator wants to predict days from a vehicle’s scheduled inspection to a battery replacement. The origin is the inspection timestamp, not vehicle purchase or the day a data export was run. Define the event as a confirmed replacement work order, including how canceled orders and repeat replacements are treated. An inspection after replacement begins a new episode only if the business question permits it.

Record what is known

For each eligible episode, store elapsed follow-up and whether replacement occurred by the last reliable observation. A vehicle still operating on day 71 is right-censored at day 71; its replacement time is unknown. Coding it as a day-71 failure biases early risk upward. Coding it as “never fails” removes the information that it survived 71 days. Right-censoring gives the statistical contract.

Distinguish observation loss

Administrative cutoff at a common export date differs from losing access when a truck changes contractor. The latter may correlate with condition or mileage, violating simple independent-censoring assumptions. Track censor reason and compare its frequency by depot, vehicle age and maintenance status before model training.

Freeze prediction-time inputs

A repair note created after battery replacement cannot predict replacement from inspection day. Snapshot odometer, voltage, route and service history as known at the inspection. Data pipelines must join by both entity and availability time. Feature availability prevents a clean-looking but unusable model.

Declare supported horizons

If very few vehicles remain observed at 180 days, a 180-day risk claim is weak even when a model returns a number. Record at-risk counts and censoring by horizon. Risk-set construction is the next step; the final target cannot be inferred from a binary label alone.

Implementation

python
inspection_episodes = [
    {"vehicle": "fleet-47", "inspection_day": 24, "event_day": 81, "last_seen_day": 92},
    {"vehicle": "fleet-62", "inspection_day": 31, "event_day": None, "last_seen_day": 102},
    {"vehicle": "fleet-83", "inspection_day": 39, "event_day": 67, "last_seen_day": 86},
]

def survival_targets(episodes):
    targets = []
    for episode in episodes:
        origin = episode["inspection_day"]
        event = episode["event_day"]
        observed = event is not None and event <= episode["last_seen_day"]
        end = event if observed else episode["last_seen_day"]
        if end < origin:
            raise ValueError("follow-up ends before inspection")
        targets.append((episode["vehicle"], end - origin, observed))
    return targets

assert survival_targets(inspection_episodes) == [
    ("fleet-47", 57, True), ("fleet-62", 71, False), ("fleet-83", 28, True)]

Performance and operating cost

Building N targets costs O(N) time and O(N) output storage. The harder expense is reliable event reconciliation and censor-reason auditing. A binary classification shortcut can seem cheaper, but it discards varying follow-up and may train on immature negatives.

Common Mistakes

  • Do not replace a censored time with an event time.
  • Do not mix purchase-date and inspection-date origins.
  • Do not include post-inspection repair fields among baseline predictors.

Read next

ai-data
machine-learning
Storage details