Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Survival-curve uncertainty and thin-tail support

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A survival estimate needs both an uncertainty interval and the number still at risk; late steps based on a few cases carry little precision.

Read beyond the line

The unresolved-case curve can look stable after day five even when only two cases remain under observation. A fixed-horizon estimate should report its event table, risk count and uncertainty. For a Kaplan–Meier step with d resolutions among n at risk, Greenwood accumulates d divided by n times (n minus d) across event times while the estimated unresolved share stays strictly between zero and one. The curve supplies the point estimate; the variance calculation describes sampling uncertainty under its assumptions.

Use bounded interval construction

A plain symmetric interval around a probability can extend below zero or above one. A log-log transformation keeps the resulting bounds in the probability range when the estimate lies strictly between zero and one. The code below returns an approximate pointwise interval at each supported event time. It deliberately leaves S equal to zero or one without this transformed interval rather than dividing by an undefined logarithm. An alternative interval method is needed at those boundaries.

Explain what the interval covers

A 95% pointwise interval describes uncertainty at a specified time under repeated samples and the estimator assumptions; it is not a 95% band that covers the entire curve at every time simultaneously. Inspecting twenty horizons and reporting the narrowest or most favorable one defeats the stated decision horizon. The interval also does not repair informative censoring, shifted status codes or a queue cohort that excludes early resolutions. The censoring contract comes first.

Show how precision collapses

In the eight-case teaching ledger, the day-two resolution has eight at risk, while the day-ten resolution has only two. The later step changes the estimated unresolved share from 0.40 to 0.20, but its Greenwood contribution is much larger. A dashboard that draws both steps with identical visual authority hides the thinning denominator. Pair the curve with at-risk counts at planned milestones and stop the comparison when a group lacks adequate follow-up.

Plan the uncertainty unit

If a customer contributes multiple cases, case-level independence may fail. Either select one case per customer, use a customer-level resampling plan, or state why the case is the independent unit. Stratified sampling and different entry dates need their own variance treatment. This simple interval is a teaching calculation for independent cases, not a substitute for a full design-aware analysis. A queue comparison adds a common horizon and a difference scale.

Implementation

python
from math import exp, log, sqrt

def unresolved_intervals(event_rows, z_score=1.96):
    unresolved = 1.0
    greenwood = 0.0
    output = []
    for day, at_risk, resolved, censored in event_rows:
        if at_risk <= 0 or min(resolved, censored) < 0:
            raise ValueError("invalid event table")
        if resolved + censored > at_risk:
            raise ValueError("more outcomes than cases at risk")
        unresolved *= 1 - resolved / at_risk
        if resolved and at_risk > resolved:
            greenwood += resolved / (at_risk * (at_risk - resolved))
        interval = None
        if 0 < unresolved < 1:
            transformed = log(-log(unresolved))
            standard_error = sqrt(greenwood) / abs(log(unresolved))
            lower = exp(-exp(transformed + z_score * standard_error))
            upper = exp(-exp(transformed - z_score * standard_error))
            interval = (lower, upper)
        output.append((day, unresolved, interval))
    return output

ledger = [(2, 8, 1, 0), (3, 7, 1, 1), (5, 5, 1, 1),
          (8, 3, 1, 0), (10, 2, 1, 1)]
estimates = unresolved_intervals(ledger)
assert round(estimates[2][1], 2) == 0.60
assert all(bounds is None or bounds[0] <= share <= bounds[1]
           for _, share, bounds in estimates)

Performance and operating cost

An already ordered U-row event table takes O(U) time and O(U) output space; reporting one chosen horizon uses O(1) additional space. Pointwise intervals become unstable in thin tails even though the loop remains cheap. Repeated customer-level bootstrap samples cost substantially more.

Common Mistakes

  • Do not present a pointwise interval as a simultaneous curve band.
  • Do not hide the day-ten risk count behind a smooth-looking plot.
  • Do not interpret the interval as protection against informative censoring.

Read next

Continue the workflow: Diagnose the proportional-hazards assumption.

Continue the workflow: Restricted mean time: compare curves at a supported horizon.

ai-data
data-science
Storage details