Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Competing events: calculate the probability of each first outcome

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Cumulative incidence tracks the probability of a particular first event while retaining other first events as competing outcomes.

Define what stops the target event

A support queue can resolve a ticket, lose it to customer withdrawal, or remain open at extraction. If withdrawal makes resolution of that ticket impossible in the defined episode, withdrawal competes with resolution. An ordinary censored record is different: its future event is unknown after follow-up stops. Define entry, event time, event type, and administrative censoring before calculating a curve. The risk-set lesson supplies the starting denominator and event order.

Accumulate target events from the full risk set

At each event time, the increment in target cumulative incidence is the probability of having avoided all first events before that time, multiplied by target events divided by the current risk set. Then reduce event-free survival by both target and competing events. The code accepts grouped counts already ordered by event time. It assumes any withdrawals from follow-up have already been removed from later risk sets; it does not infer their timing. A later target event cannot be assigned to a ticket already withdrawn in this single-episode analysis.

Avoid the wrong censoring shortcut

If competing withdrawals are treated like ordinary independent censoring and one minus a target-only Kaplan–Meier curve is called the real-world resolution probability, the result generally overstates that probability. That shortcut imagines a setting where withdrawn tickets could still later resolve. It may have a different hypothetical interpretation, but it is not the observed-world chance of resolution before withdrawal. The next lesson distinguishes a conditional event rate from an absolute event probability.

Show every state at the horizon

At each decision horizon, report the chance of resolution, the chance of withdrawal, and the chance of no first event yet. They should fit the same episode definition and sum to one before any loss-to-follow-up adjustment. State uncertainty and how censoring was handled. A short study with late entry or many open tickets cannot be summarized by dividing resolved tickets by all tickets without regard to follow-up. The project makes the event ledger a release gate.

Implementation

python
def grouped_cumulative_incidence(event_rows):
    survival = 1.0
    target_cumulative = 0.0
    other_cumulative = 0.0
    results = []
    for at_risk, resolved, withdrawn in event_rows:
        if at_risk <= 0 or min(resolved, withdrawn) < 0 or            resolved + withdrawn > at_risk:
            raise ValueError("invalid event-time counts")
        target_cumulative += survival * resolved / at_risk
        other_cumulative += survival * withdrawn / at_risk
        survival *= 1 - (resolved + withdrawn) / at_risk
        results.append((target_cumulative, other_cumulative, survival))
    return results

curve = grouped_cumulative_incidence([(50, 4, 6), (40, 5, 3)])
assert tuple(round(value, 2) for value in curve[-1]) == (0.18, 0.18, 0.64)

Performance and operating cost

The grouped scan takes O(k) time and O(k) output space for k event times; a final-horizon value needs O(1) extra space. Sorting raw events costs O(n log n). The hard part is reconciling event types and risk sets when tickets re-open or follow-up ends at different times.

Common Mistakes

  • Treating withdrawal as independent censoring when reporting observed-world resolution probability.
  • Letting a ticket resolve after it has left the single-episode risk set.
  • Mixing administrative censoring with a real competing outcome.
  • Reporting a cumulative incidence without its time horizon or remaining event-free fraction.

Read next

ai-data
applied-statistics
Storage details