A proportional-hazards model assigns one time-constant covariate effect; systematic changes in event-time residuals can expose a misleading single hazard ratio.
Diagnose the proportional-hazards assumption
Ask whether one ratio is coherent
A priority workflow may accelerate early resolutions but have little effect on difficult cases still open after two weeks. One constant priority hazard ratio then compresses distinct phases into a number that may answer neither phase well. Compare group curves, event counts and at-risk counts over elapsed time before fitting. A crossing or changing separation warrants closer examination, although visual patterns alone are not a formal test. The unresolved curves establish the observed shape.
Inspect event-time residuals
For a binary priority covariate, a Schoenfeld-style residual at a resolution event is the event case’s priority indicator minus its risk-set-weighted expected priority indicator under the fitted coefficient. If residuals have a systematic relationship with event time, the assumed constant effect may be doubtful. The code computes raw residuals for untied events; it does not standardize them, estimate their variance, or produce a formal significance test. A production review uses a supported residual diagnostic and plots alongside the business timeline.
Avoid mechanical p-value decisions
With many covariates, some checks flag by chance. With a very large cohort, tiny departures can yield small p-values even when fixed-horizon decisions barely change. Review effect shape, late risk-set support, and whether the difference matters to the intended policy. A failed diagnostic does not identify why the relationship changed; the issue might be genuinely time-dependent treatment, shifting case mix, delayed entry, or misrecorded resolution times.
Choose a revised estimand or model
If priority impact differs early and late, report time-specific contrasts or a prespecified restricted-mean difference at a common horizon. A stratified Cox fit can handle a nonproportional nuisance category when its coefficient is not the desired output. A time-varying coefficient model needs an explicit interval structure and adequate events in each period. Do not search many cut points after seeing the residual plot and present the best-looking split as a prespecified result. Bounded open-time comparison avoids requiring one constant hazard ratio.
Keep causal claims separate
A proportional-hazards diagnostic concerns model form, not confounding. Even a perfectly flat residual pattern cannot make queue assignment random. Record priority rules, pre-assignment severity and policy version alongside the diagnostic. If the question is whether prioritization causes faster resolution, build a design that addresses treatment selection. The model-review project treats this as a release gate, not a decorative chart.
Implementation
from math import exp
def priority_event_residuals(case_records, coefficient):
event_days = [day for _, day, resolved, _ in case_records if resolved]
if len(event_days) != len(set(event_days)):
raise ValueError("untied event times required")
residuals = []
for _, event_day, resolved, priority in case_records:
if not resolved:
continue
risk_set = [(peer_priority, exp(coefficient * peer_priority))
for _, peer_day, _, peer_priority in case_records
if peer_day >= event_day]
total_weight = sum(weight for _, weight in risk_set)
expected_priority = sum(indicator * weight for indicator, weight in risk_set) / total_weight
residuals.append((event_day, priority - expected_priority))
return sorted(residuals)
cases = [("P61", 2, 1, 1), ("P62", 3, 1, 0), ("P63", 4, 0, 0),
("P64", 5, 1, 1), ("P65", 6, 0, 1), ("P66", 7, 1, 0),
("P67", 8, 1, 1), ("P68", 9, 0, 0)]
residuals = priority_event_residuals(cases, 0.35)
assert len(residuals) == 5
assert all(-1 <= value <= 1 for _, value in residuals)
assert [day for day, _ in residuals] == [2, 3, 5, 7, 8]Performance and operating cost
The direct implementation scans N cases for each of E events: O(E × N) time and O(E) output space. It is a transparent raw residual calculation. Scaling, uncertainty tests, tied events and clustered cases require more machinery than this exercise includes.
Common Mistakes
- Do not call an unscaled residual scan a formal assumption test.
- Do not treat a flat residual plot as proof of no confounding.
- Do not choose time splits only after inspecting the pattern.
