Discrimination, calibration and decision value must be measured at horizons with enough observed follow-up; ordinary binary metrics cannot treat all censored cases as negatives.
Survival-model evaluation at supported horizons
Name the output
An operations dashboard may need the probability of battery replacement within 45 days of inspection. A hazard ratio is not this probability, and a risk ranking alone does not supply a calibrated 45-day forecast. Convert the fitted model to a survival or cumulative-risk estimate under its assumptions, then validate at the specified horizon.
Handle unknown horizon outcomes
For a 45-day test, a vehicle replaced on day 23 is an event, and one observed event-free through day 45 is a known nonevent. A vehicle last seen event-free on day 29 has unknown day-45 status. Dropping it naively may bias results if censoring is selective. Use a justified censoring-aware estimator and test support before claiming calibration. Censor-weighted Brier scoring gives one approach.
Separate ranking from probabilities
Concordance summarizes ordering among comparable pairs, while a horizon Brier score and calibration bins assess predicted probabilities. A model may rank older batteries correctly yet consistently overstate 45-day failures. Report both aspects, plus event counts and uncertainty by depot. Fixed-horizon calibration expands the diagnostic.
Inspect the censor process
Estimate observation survival by route, depot and maintenance status. Very large inverse-censor weights signal weak support; cap the claimed horizon or gather more follow-up rather than trusting a high-variance score. Do not tune the censor model or decision threshold on the final test.
Tie metrics to action
If only 26 vehicle inspections can be prioritized weekly, compare the number of replacements captured in the highest-risk eligible cohort, with a clear handling rule for censored outcomes. The dashboard should show uncertainty and avoid pretending that a thin 180-day tail is as reliable as the 45-day estimate.
Implementation
episodes = [
{"vehicle": "fleet-47", "follow_up": 57, "event": True, "risk_45": 0.24},
{"vehicle": "fleet-62", "follow_up": 71, "event": False, "risk_45": 0.08},
{"vehicle": "fleet-83", "follow_up": 28, "event": True, "risk_45": 0.43},
{"vehicle": "fleet-94", "follow_up": 19, "event": False, "risk_45": 0.12},
]
def known_horizon_status(record, horizon):
if record["event"] and record["follow_up"] <= horizon:
return 1
if record["follow_up"] >= horizon:
return 0
return None
statuses = {row["vehicle"]: known_horizon_status(row, 45) for row in episodes}
assert statuses == {"fleet-47": 0, "fleet-62": 0,
"fleet-83": 1, "fleet-94": None}Performance and operating cost
Assigning known horizon status costs O(N) time and O(N) output storage. The code intentionally leaves early-censored outcomes unknown; it is not a censoring-adjusted score. Censoring estimation, resampling for uncertainty and full survival-curve prediction add compute and data requirements.
Common Mistakes
- Do not count early-censored cases as day-45 nonevents.
- Do not equate a hazard ratio with an absolute failure probability.
- Do not report a distant-horizon score without at-risk and censor support.
