Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Landmark analysis and immortal-time traps

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A post-entry action must not be treated as if it were known at case creation; doing so gives its future recipients guaranteed event-free time before the action.

Locate the classification error

Suppose escalation may happen on day three or day six after a case opens. Labelling every eventually escalated case as escalated from day zero means a case must remain unresolved long enough to earn that label. A case resolved on day two could never enter the escalated group. The apparent early advantage is built into the grouping rule. This is an immortal-time error, not a subtle model-tuning problem. Entry and selection should be checked before drawing curves.

Choose a landmark and reset the question

At a prespecified day-four landmark, retain only cases still unresolved and observed immediately after day four. Classify exposure using escalation history strictly before that instant, then follow both groups from the landmark. A case escalated on day six is unexposed at day four, even if it escalates later. The estimand is now conditional on surviving unresolved to day four; do not report it as the effect among all cases at creation.

Handle timestamp boundaries

If escalation and resolution share a rounded day-four timestamp, the order is unknown. Reject or adjudicate the record using finer timestamps rather than assigning whichever order helps the analysis. A censor at day three is absent from the day-four landmark cohort, while a censor at day seven is eligible at day four. Keep the chosen boundary convention in the extraction code and in the report. Risk sets depend on these choices.

Know what the design cannot prove

Escalation is usually triggered by case difficulty, so landmark grouping is still confounded by severity and prior agent action. A lower resolution rate after escalation does not establish that escalation caused delay. Adjust for defensible pre-landmark attributes or use a designed intervention if the goal is causal. A time-varying exposure model can use later status changes without freezing them at day zero, but it still needs careful confounding assumptions.

Write down the eligible population

The code emits one row per eligible case with its day-four exposure state and residual follow-up. It excludes a day-two resolution, keeps a case escalated at day three, and keeps a case escalating at day six in the day-four unexposed group. A second landmark would define a different population and should not be selected after inspecting outcomes. The audit project requires an explicit decision about whether this conditional view is even needed.

Implementation

python
def landmark_cases(case_records, landmark_day):
    if landmark_day <= 0:
        raise ValueError("landmark must be positive")
    if len({case_id for case_id, _, _, _ in case_records}) != len(case_records):
        raise ValueError("duplicate case ID")
    eligible = []
    for case_id, exit_day, resolved, escalation_day in case_records:
        if exit_day < 0 or resolved not in (0, 1):
            raise ValueError("invalid outcome")
        if escalation_day is not None:
            if escalation_day < 0 or escalation_day > exit_day:
                raise ValueError("impossible escalation time")
            if escalation_day == landmark_day and exit_day >= landmark_day:
                raise ValueError("ambiguous landmark ordering")
        if exit_day <= landmark_day:
            continue
        was_escalated = escalation_day is not None and escalation_day < landmark_day
        eligible.append((case_id, was_escalated, exit_day - landmark_day, resolved))
    return eligible

cases = [("L41", 2, 1, None), ("L42", 8, 1, 3),
         ("L43", 10, 0, 6), ("L44", 7, 1, None)]
cohort = landmark_cases(cases, 4)
assert cohort == [("L42", True, 4, 1), ("L43", False, 6, 0),
                  ("L44", False, 3, 1)]

Performance and operating cost

Selecting a landmark cohort from N records costs O(N) time and O(N) output space. Repeating many candidate landmarks can be expensive and invites selective reporting. The larger cost is interpreting a conditional observational comparison as a causal intervention.

Common Mistakes

  • Do not assign eventual escalation status to the case at day zero.
  • Do not include a case that resolved before the landmark.
  • Do not call a landmark association a causal effect without a defensible design.

Read next

Continue the workflow: Time-updated case histories and the observation clock.

ai-data
data-science
Storage details