A difference in restricted mean open time compares the area under two unresolved-case curves over the same prespecified horizon, in case-days.
Compare queues by restricted mean open time
Keep the contrast in its own units
If queue A has a six-day restricted mean of 4.975 open case-days and queue B has 4.250, the A-minus-B difference is 0.725 case-days through day six. It is not a 0.725-day difference in eventual mean resolution time. Nor does it prove a routing change caused the difference. The two cohorts must share a time origin, outcome definition, calendar eligibility window, censoring policy and horizon. The single-curve calculation defines the area.
Check common support
A queue introduced only last week may have many cases censored before day six. Its day-six estimate has little support even if a function returns a number. Record at-risk counts at the horizon and a minimum acceptable follow-up rule before comparison. If one queue lacks support, shorten the common horizon for both or wait for more mature data. Never let each queue choose its own maximum follow-up: the resulting areas answer different questions. Tail diagnostics make this visible.
Use a worked contrast without hidden censoring
For a clean arithmetic check, six A cases end at days two, four, five, eight, eight and eight; the final three remain unresolved beyond the day-six horizon. Six B cases end at days one, two, four, five, eight and eight; the final two also remain unresolved beyond day six. With no censoring before day six, area under each unresolved curve equals the average of each duration capped at six. A is 29/6, or about 4.833 case-days; B is 24/6, or four. The difference is 5/6 case-day. This fixture checks code, not a population claim.
Add uncertainty at the correct unit
For a real comparison, recompute both curves and their area difference in each bootstrap replicate, resampling independent customers within queue if one customer can open several cases. Retain cohort and censoring structure within replicates. A narrow interval from resampling individual status updates would be false precision. If queue assignment depends on case complexity, uncertainty about the descriptive difference is separate from bias in a causal claim. Bootstrap units need an explicit contract.
Pair open time with exit types
A queue can show shorter open time because it resolves cases faster or because it cancels more of them. Publish resolution and cancellation cumulative incidence at the same horizon beside the bounded open-time difference. For the first metric, state whether the curve counts only resolution or any terminal exit; changing that rule changes the area. Competing outcomes prevent a fast cancellation from masquerading as better resolution.
Implementation
def capped_open_time(case_observations, horizon_days):
if horizon_days <= 0 or not case_observations:
raise ValueError("missing cases or horizon")
if any(duration < horizon_days for duration, observed_event in case_observations
if not observed_event):
raise ValueError("early censoring needs a Kaplan-Meier area estimate")
if any(duration < 0 or observed_event not in (0, 1)
for duration, observed_event in case_observations):
raise ValueError("invalid observation")
return sum(min(duration, horizon_days) for duration, _ in case_observations) / len(case_observations)
queue_a = [(2, 1), (4, 1), (5, 1), (8, 0), (8, 0), (8, 0)]
queue_b = [(1, 1), (2, 1), (4, 1), (5, 1), (8, 0), (8, 0)]
queue_a_days = capped_open_time(queue_a, 6)
queue_b_days = capped_open_time(queue_b, 6)
assert queue_a_days == 29 / 6
assert queue_b_days == 4
assert abs((queue_a_days - queue_b_days) - 5 / 6) < 1e-12Performance and operating cost
For N case records without censoring before the selected horizon, direct capped averaging is O(N) time and O(1) additional space. With early censoring, compute separate Kaplan–Meier curves at O(N log N) sorting cost, then integrate. Resampling B replicates multiplies that work by B.
Common Mistakes
- Do not apply capped averaging when a case was censored before the horizon.
- Do not compare queues over different horizons.
- Do not mistake a descriptive queue difference for an effect of queue assignment.
