Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Aalen–Johansen cumulative incidence from case events

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A cumulative-incidence curve accumulates one event type using the probability of remaining free from every terminal event just before each event time.

Start with the event-free pool

Suppose five support cases are followed from creation. At each distinct terminal time, the risk set contains cases still free of resolution and cancellation immediately beforehand. An observed resolution adds the current event-free share divided by the risk-set size to resolution incidence; a cancellation adds the same type of increment to cancellation incidence. Both terminal types reduce the future event-free share. Right censoring only removes a case from later risk sets. Risk sets determine the denominator.

Inspect the ledger calculation

The fixture below has resolution at days two and seven, cancellation at day three, and status loss at days five and nine. Its final resolution incidence is 0.50, cancellation incidence 0.20 and event-free share 0.30. These are small-fixture arithmetic checks, not reliable population estimates. In a real cohort, show uncertainty and the number still under observation near each plotted time. A flat late curve can mean no events or no cases left.

Keep the probability identity

At each time, estimated resolution incidence plus cancellation incidence plus event-free probability should be one, allowing tiny floating-point error. If two categories overlap, the identity may fail or a case may leave twice. An event after the extraction cutoff must not enter an earlier curve. A case observed open at the cutoff is censored there; it is not a third terminal event.

Use a supported tie method

The teaching routine rejects shared event ages so the update order is unambiguous. Actual systems often record statuses only to the day, creating ties. A maintained estimator groups events by time and handles all event types in the group together. Do not jitter dates merely to force a preferred ordering. The observation interval may be the meaningful unit, in which case interval-censored methods and sensitivity analysis are needed. Interval censoring is a different information pattern.

Interpret incidence as observed-world risk

The curve estimates the probability of experiencing a particular first event by a horizon in a population with the observed competing exits, subject to censoring assumptions. It does not give the hazard among cases still open, nor the probability of resolution if all cancellations vanished. Compare workflows on a prespecified arrival cohort and horizon, then check how their competing-event policies differ.

Implementation

python
def competing_incidence(case_records):
    if len({case_id for case_id, _, _ in case_records}) != len(case_records):
        raise ValueError("duplicate case ID")
    if any(day <= 0 or outcome not in {"resolved", "cancelled", "censored"}
           for _, day, outcome in case_records):
        raise ValueError("invalid case observation")
    if len({day for _, day, _ in case_records}) != len(case_records):
        raise ValueError("teaching calculation requires distinct times")
    at_risk = len(case_records)
    event_free = 1.0
    incidence = {"resolved": 0.0, "cancelled": 0.0}
    steps = []
    for _, day, outcome in sorted(case_records, key=lambda record: record[1]):
        if outcome != "censored":
            increment = event_free / at_risk
            incidence[outcome] += increment
            event_free *= 1 - 1 / at_risk
        steps.append((day, incidence["resolved"], incidence["cancelled"],
                      event_free, at_risk))
        at_risk -= 1
    return steps

case_events = [("E311", 2, "resolved"), ("E312", 3, "cancelled"),
               ("E313", 5, "censored"), ("E314", 7, "resolved"),
               ("E315", 9, "censored")]
curve = competing_incidence(case_events)
assert all(abs(actual - expected) < 1e-12
           for actual, expected in zip(curve[-1][1:4], (0.5, 0.2, 0.3)))
assert all(abs(resolved + cancelled + open_share - 1) < 1e-12
           for _, resolved, cancelled, open_share, _ in curve)

Performance and operating cost

Sorting N case records is O(N log N) time and the step curve uses O(N) space. The per-time update is O(1) for two terminal causes. Production estimators add grouped ties, variance calculation and checks for censoring support.

Common Mistakes

  • Do not remove a competing event as if it were ordinary loss of follow-up.
  • Do not infer a stable late probability from a tiny risk set.
  • Do not create artificial within-day ordering to evade tied times.

Read next

ai-data
data-science
Storage details