Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Counting-process intervals for a time-varying Cox model

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A start-stop survival record assigns the covariate state that applied during each open interval and places an event only at its terminal stop.

One case can occupy several rows

A case open from hour zero to twelve and escalated at hour five contributes (0,5] with priority zero and (5,12] with priority one if resolution occurs at twelve. The case ID stays the same. The rows partition the observed lifetime without gaps or overlap, and only the last row carries a resolution event. A censor at twelve ends the final row without an event. The observation-clock lesson decides whether the escalation was actually available at hour five.

Read the event-time risk set correctly

At event time t, a row is active when start is strictly below t and stop is at least t. The event row receives the state applicable just before its stop. At a shared boundary, the earlier interval ends and the later one begins, so an update recorded exactly at an event time needs an explicit ordering rule. The teaching code rejects coincident update and terminal times to keep that rule visible rather than guessing.

Build an audit before fitting

Verify one terminal record per case, increasing boundaries, no zero-length rows, no update after terminal observation and no event on an intermediate row. A static Cox model fed duplicated interval rows as independent cases has the wrong risk set and uncertainty. Use a fitter that accepts case IDs, start, stop, event and time-varying covariates; record its handling of tied events and variance. Static partial likelihood explains the risk-set numerator and denominator.

Keep prediction and causal questions apart

An escalation can be triggered by a case becoming difficult. Its fitted hazard association is not the effect of forcing escalation. Time-updated severity may also be affected by earlier priority actions; ordinary adjustment can then bias a causal comparison. For prediction, later state is usable only at or after its observed update, and future state is unknown. For intervention effects, specify a separate design and assignment mechanism.

Inspect intervals at boundaries

For each event hour, count unique active case IDs and compare them with the raw case ledger. A case must never occupy two active rows at the same event. Check how many events occur after the last update in each state. Sparse late risk sets make a time-varying coefficient fragile even when software reports a number. Effect-shape diagnostics remain necessary after interval conversion.

Implementation

python
def case_intervals(case_id, terminal_hour, resolved, initial_priority, updates):
    if terminal_hour <= 0 or resolved not in (0, 1) or initial_priority not in (0, 1):
        raise ValueError("invalid terminal record")
    if len({hour for hour, _ in updates}) != len(updates):
        raise ValueError("duplicate update hour")
    boundaries = [(0, initial_priority)] + sorted(updates)
    if any(hour <= 0 or hour >= terminal_hour or state not in (0, 1)
           for hour, state in updates):
        raise ValueError("update outside observed lifetime")
    result = []
    for position, (start_hour, priority_state) in enumerate(boundaries):
        stop_hour = (boundaries[position + 1][0]
                     if position + 1 < len(boundaries) else terminal_hour)
        result.append((case_id, start_hour, stop_hour, priority_state,
                       int(bool(resolved and stop_hour == terminal_hour))))
    return result

history = case_intervals("R317", 12, 1, 0, [(5, 1)])
assert history == [("R317", 0, 5, 0, 0), ("R317", 5, 12, 1, 1)]
assert sum(start < 9 <= stop for _, start, stop, _, _ in history) == 1

Performance and operating cost

Sorting U updates costs O(U log U) time and creates U + 1 intervals, so output space is O(U). Across N cases, a naive event-time risk-set scan can be quadratic in total rows; indexed survival fitters avoid repeatedly rebuilding every set. The code prepares records, not a coefficient estimate.

Common Mistakes

  • Do not mark an intermediate interval as the event.
  • Do not enter an update at or after the terminal event as if it preceded that event.
  • Do not interpret a time-updated association as a randomized treatment effect.

Read next

ai-data
data-science
Storage details