A censor-aware Brier score compares predicted survival at a fixed horizon with observed event status, weighting cases whose outcome is known by their chance of remaining observable.
IPCW Brier score for censored survival predictions
Freeze prediction and horizon
Suppose a case-created model predicts the probability of remaining unresolved through hour nine. Keep that number as it was at case creation. For a case resolved at hour six, the nine-hour survival outcome is zero; for a case still observed open at hour eleven, it is one. A case lost at hour five has unknown nine-hour status and contributes no direct squared error. Fully observed calibration is simpler when every status is known.
Use two weight locations
The score weights an event at t at or before nine by inverse G(t-) and a case observed beyond nine by inverse G(9). G is censoring survival estimated on a suitable training cohort, not fitted by peeking at a held-out test outcome. The code supplies a tiny fixed censoring curve to expose the arithmetic. It divides the weighted sum by the full test cohort size; dropping censored rows and dividing by the remaining count changes the estimand.
Separate score from ranking
A low Brier score rewards useful absolute survival probabilities. A model that gives every case the same middle probability can still rank no one, and a strong ranker can have biased probabilities. Compare against a prespecified reference such as the training cohort survival curve, and inspect calibration at several supported horizons. Do not choose a single hour after reviewing the score curve and present it as planned.
Inspect censor-model support
Before calculating, verify G(event time-) and G(h) are positive. If the minimum supported value is tiny, extreme weights can make one record dominate. State the score horizon, number at risk, censoring fraction, weight range and uncertainty. The teaching fixture has distinct event and censor times. Real data require a method with documented tie conventions and confidence estimates. Censoring weights explain those prerequisites.
Know what this score does not prove
An IPCW score does not certify that the model transfers to a newer workflow or that censoring is independent after conditioning. A model evaluated on the training month may benefit from changed queue rules or retrospective notes. Lock the preprocessing, fit and censoring estimator before looking at later cases. Temporal validation checks that transfer.
Implementation
def step_value(steps, target_hour, just_before=False):
value = 1.0
for step_hour, before, after in steps:
if step_hour < target_hour or (step_hour == target_hour and not just_before):
value = after
else:
break
return value
def ipcw_brier_at_horizon(predictions, censor_steps, horizon_hour):
if horizon_hour <= 0 or not predictions:
raise ValueError("invalid horizon or empty cohort")
if len({case_id for case_id, _, _, _ in predictions}) != len(predictions):
raise ValueError("duplicate case ID")
weighted_error = 0.0
for case_id, survival_probability, observed_hour, resolved in predictions:
if not 0 <= survival_probability <= 1 or observed_hour <= 0:
raise ValueError("invalid prediction")
if resolved not in (0, 1):
raise ValueError("invalid outcome flag")
if resolved and observed_hour <= horizon_hour:
observation_chance = step_value(censor_steps, observed_hour, True)
outcome = 0
elif observed_hour > horizon_hour:
observation_chance = step_value(censor_steps, horizon_hour)
outcome = 1
else:
continue
if observation_chance <= 0:
raise ValueError("unsupported horizon or event time")
weighted_error += (outcome - survival_probability) ** 2 / observation_chance
return weighted_error / len(predictions)
training_censor_curve = [(4, 1.0, 0.75), (8, 0.75, 0.5)]
test_predictions = [("R501", .18, 3, 1), ("R502", .72, 12, 0),
("R503", .55, 5, 0), ("R504", .31, 7, 1)]
score = ipcw_brier_at_horizon(test_predictions, training_censor_curve, 9)
assert abs(score - (0.0324 + 0.0961 / 0.75 + 0.0784 / 0.5) / 4) < 1e-10Performance and operating cost
With N predictions and C censoring steps, the direct lookup is O(N × C) time and O(1) extra space beyond inputs. Binary search on sorted step times lowers lookup to O(N log C). Statistical support matters more than lookup speed: near-zero G produces high-variance weights.
Common Mistakes
- Do not label a case censored before the horizon as unresolved at the horizon.
- Do not divide only by complete cases after using IPCW.
- Do not estimate censoring from a held-out cohort and then describe its evaluation as untouched.
Read next
- Censoring survival weights and horizon support
- Calibrate resolution predictions at a fixed horizon
- Temporal validation of a support-resolution model
- Project: validate a time-updated resolution model
Continue the workflow: Paired bootstrap uncertainty for prediction-score differences.
