Competing-risk calibration compares predicted first-event incidence with the observed share of that event by a fixed horizon on the same starting cohort.
Calibrate competing-risk predictions at a fixed horizon
Score the event of interest
For nine-day resolution prediction, a case resolved by day nine has target one. A case cancelled first by day nine has target zero: its resolution status is known, not missing. A case still observed open through day nine is also zero. A cancellation after day nine means the case had not resolved by the horizon, so the target remains zero. Early loss of observation is unknown and needs a censor-aware estimator rather than a forced label. The outcome contract sets this rule.
Keep the full probability vector
A model may issue resolution and cancellation probabilities for the same horizon. Check that both are in [0,1] and their sum is no more than one; the remainder is predicted still-event-free probability. The example groups cases by predicted resolution risk and compares mean prediction with observed resolution share. It also computes the mean squared error of resolution risk against the binary outcome, a fixed-horizon Brier score on this fully observed fixture.
Do not confuse calibration with causation
If a high-risk group predicts 0.68 resolution and 0.50 actually resolve, the absolute predictions overstate this event on that group. The result does not say an intervention caused the gap. Check group size, case mix, censoring, outcome coding and time period first. A small group can show a large apparent gap by chance. Use prespecified bins or a supported smooth estimator with uncertainty; do not tune bins until the chart looks attractive.
Handle incomplete follow-up correctly
The code rejects a case lost before the horizon. In a real validation cohort, estimate incidence in risk groups with a competing-risk method and use a censor-aware score with an appropriate observation model. A complete-case-only average may change the population if disappearing cases differ by severity. Censoring support sets a limit on the longest useful horizon.
Compare a frozen reference
A forecast is useful only relative to the decision it supports and a stable alternative. Compare its Brier score with a prespecified training-cohort incidence forecast on the same held-out cases; show calibration for resolution and cancellation separately. If one event policy changed, a good score in the old period may not transfer. Paired uncertainty prevents a tiny score difference from being overstated.
Implementation
def resolution_calibration(predictions, horizon_days):
if horizon_days <= 0 or not predictions:
raise ValueError("invalid horizon or empty cohort")
if len({case_id for case_id, _, _, _, _ in predictions}) != len(predictions):
raise ValueError("duplicate case ID")
groups = {"lower": [], "higher": []}
squared_errors = []
for case_id, resolution_risk, cancellation_risk, observed_day, outcome in predictions:
if (not 0 <= resolution_risk <= 1 or
not 0 <= cancellation_risk <= 1 or
resolution_risk + cancellation_risk > 1 or observed_day < 0):
raise ValueError("invalid probability or observation")
if outcome not in {"resolved", "cancelled", "open", "lost"}:
raise ValueError("unknown outcome")
if outcome in {"open", "lost"} and observed_day < horizon_days:
raise ValueError("outcome unknown at horizon")
target = int(outcome == "resolved" and observed_day <= horizon_days)
group = "higher" if resolution_risk >= .5 else "lower"
groups[group].append((resolution_risk, target))
squared_errors.append((resolution_risk - target) ** 2)
summary = {name: (len(values),
sum(risk for risk, _ in values) / len(values),
sum(target for _, target in values) / len(values))
for name, values in groups.items() if values}
return summary, sum(squared_errors) / len(squared_errors)
cases = [("E441", .72, .12, 3, "resolved"),
("E442", .61, .25, 5, "cancelled"),
("E443", .28, .37, 11, "open"),
("E444", .31, .15, 7, "resolved")]
groups, score = resolution_calibration(cases, 9)
assert groups["higher"][0] == 2
assert groups["higher"][2] == .5
assert 0 < score < 1Performance and operating cost
For N fully observed cases and a fixed number of groups, calculation is O(N) time and O(N) space because individual errors and bin members are retained. Running totals can make extra space O(1). Censor-aware scoring requires an additional observation model and support checks.
Common Mistakes
- Do not label cancellation as missing for observed-world resolution incidence.
- Do not score early-loss cases as known negatives.
- Do not claim calibration from a ranking metric alone.
