Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Survival-model evaluation at supported horizons

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Discrimination, calibration and decision value must be measured at horizons with enough observed follow-up; ordinary binary metrics cannot treat all censored cases as negatives.

Name the output

An operations dashboard may need the probability of battery replacement within 45 days of inspection. A hazard ratio is not this probability, and a risk ranking alone does not supply a calibrated 45-day forecast. Convert the fitted model to a survival or cumulative-risk estimate under its assumptions, then validate at the specified horizon.

Handle unknown horizon outcomes

For a 45-day test, a vehicle replaced on day 23 is an event, and one observed event-free through day 45 is a known nonevent. A vehicle last seen event-free on day 29 has unknown day-45 status. Dropping it naively may bias results if censoring is selective. Use a justified censoring-aware estimator and test support before claiming calibration. Censor-weighted Brier scoring gives one approach.

Separate ranking from probabilities

Concordance summarizes ordering among comparable pairs, while a horizon Brier score and calibration bins assess predicted probabilities. A model may rank older batteries correctly yet consistently overstate 45-day failures. Report both aspects, plus event counts and uncertainty by depot. Fixed-horizon calibration expands the diagnostic.

Inspect the censor process

Estimate observation survival by route, depot and maintenance status. Very large inverse-censor weights signal weak support; cap the claimed horizon or gather more follow-up rather than trusting a high-variance score. Do not tune the censor model or decision threshold on the final test.

Tie metrics to action

If only 26 vehicle inspections can be prioritized weekly, compare the number of replacements captured in the highest-risk eligible cohort, with a clear handling rule for censored outcomes. The dashboard should show uncertainty and avoid pretending that a thin 180-day tail is as reliable as the 45-day estimate.

Implementation

python
episodes = [
    {"vehicle": "fleet-47", "follow_up": 57, "event": True, "risk_45": 0.24},
    {"vehicle": "fleet-62", "follow_up": 71, "event": False, "risk_45": 0.08},
    {"vehicle": "fleet-83", "follow_up": 28, "event": True, "risk_45": 0.43},
    {"vehicle": "fleet-94", "follow_up": 19, "event": False, "risk_45": 0.12},
]

def known_horizon_status(record, horizon):
    if record["event"] and record["follow_up"] <= horizon:
        return 1
    if record["follow_up"] >= horizon:
        return 0
    return None

statuses = {row["vehicle"]: known_horizon_status(row, 45) for row in episodes}
assert statuses == {"fleet-47": 0, "fleet-62": 0,
                    "fleet-83": 1, "fleet-94": None}

Performance and operating cost

Assigning known horizon status costs O(N) time and O(N) output storage. The code intentionally leaves early-censored outcomes unknown; it is not a censoring-adjusted score. Censoring estimation, resampling for uncertainty and full survival-curve prediction add compute and data requirements.

Common Mistakes

  • Do not count early-censored cases as day-45 nonevents.
  • Do not equate a hazard ratio with an absolute failure probability.
  • Do not report a distant-horizon score without at-risk and censor support.

Read next

ai-data
machine-learning
Storage details