Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Event-time risk sets and tied outcomes

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A risk set counts cases still observable and eligible immediately before an event time; its denominator changes as events and censoring occur.

Fix the clock before counting

For a support-resolution analysis, time zero is case creation and the event is the first recorded resolution. A case still open at the audit cutoff is right-censored at its observed age. A cancelled case needs an explicit outcome policy: it is a competing event if cancellation prevents ordinary resolution, not an ordinary unresolved case. Record timestamps, status, cohort entry and extraction cutoff together. Right-censoring explains why an open case still contributes observed time.

Build the denominator at each event time

Use eight cases with observed durations and event flags: (2,1), (3,0), (3,1), (5,1), (5,0), (8,1), (10,0), (10,1). At day two, eight cases are at risk. At day three, seven remain; one resolves and one is censored. Both belong to the day-three risk set immediately before either outcome. After processing that time, five remain. The risk set at day five is five, then three at day eight and two at day ten.

State a tie convention

At a shared timestamp, count event cases against the pre-time risk set, then remove both events and censored cases before the next time. This convention prevents a day-three censor from shrinking the denominator before the day-three resolution. If source timestamps are rounded to whole days, ties can be frequent; keep the time unit and tie rule in the analysis contract. Do not fabricate an arbitrary within-day ordering just because rows arrived in a particular file order.

Check conservation

At each distinct time, the next risk count equals the current count minus resolutions minus censored cases. The sequence here is eight, seven, five, three, two, then zero. An event after a case was censored, a negative duration, or duplicate case ID violates the input contract. Reconcile event counts with the operational status ledger before estimating a curve. Dataset grain matters: one row per case, not one row per status update.

Expose the event table

Keep columns for time, at risk, resolved, censored and remaining. That table is the audit trail for the Kaplan–Meier curve, competing-outcome calculations and later horizon summaries. When cases enter observation late because the extract includes only cases that reached a specific queue, a simple case-created cohort is no longer valid; account for delayed entry or choose a new time origin.

Implementation

python
def event_table(case_observations):
    if len({case_id for case_id, _, _ in case_observations}) != len(case_observations):
        raise ValueError("duplicate case ID")
    by_day = {}
    for case_id, duration_days, resolved in case_observations:
        if duration_days < 0 or resolved not in (0, 1):
            raise ValueError("invalid duration or event flag")
        counts = by_day.setdefault(duration_days, {"resolved": 0, "censored": 0})
        counts["resolved" if resolved else "censored"] += 1
    at_risk = len(case_observations)
    table = []
    for day, counts in sorted(by_day.items()):
        if counts["resolved"] + counts["censored"] > at_risk:
            raise ValueError("risk count exhausted")
        table.append((day, at_risk, counts["resolved"], counts["censored"]))
        at_risk -= counts["resolved"] + counts["censored"]
    return table

cases = [("C21", 2, 1), ("C22", 3, 0), ("C23", 3, 1),
         ("C24", 5, 1), ("C25", 5, 0), ("C26", 8, 1),
         ("C27", 10, 0), ("C28", 10, 1)]
assert event_table(cases) == [(2, 8, 1, 0), (3, 7, 1, 1),
                              (5, 5, 1, 1), (8, 3, 1, 0), (10, 2, 1, 1)]

Performance and operating cost

Grouping N case records is O(N) expected time and O(U) space for U distinct durations; sorting those durations costs O(U log U). A single chronological pass then costs O(U). Bad cohort definitions dominate computational cost as a source of error.

Common Mistakes

  • Do not remove same-time censors before counting same-time events.
  • Do not treat status-update rows as independent cases.
  • Do not label a cancellation as ordinary censoring without an outcome policy.

Read next

Continue the workflow: Delayed entry and left-truncated risk sets.

ai-data
data-science
Storage details