To test whether a model has lost its relation to the outcome, compare errors on mature labels within comparable input slices and time windows.
Concept change with mature outcomes
Separate a changing mix from a changing relation
A depot may receive more high-backlog shipments without a change in within-band miss rates. That is an input-mix problem. If the miss rate within a comparable band rises after a routing rule changes, the input-to-outcome relation may have changed. Neither pattern is proven by six rows; the example is a transparent diagnostic. Support overlap is the first check.
Wait for eligible labels
An unresolved shipment has no final handoff outcome yet. Comparing only fast-resolving cases can produce a false improvement. Define a maturity delay and a fixed cutoff before calculating error. Keep the original issued probability and threshold alongside the later outcome. The label-clock lesson treats unresolved cases as missing evidence, not negatives.
Compare within a declared slice
The code contrasts missed-handoff prevalence and alert error in the same middle-backlog band across two periods. Prevalence can change because of an unrecorded cause; observed differences justify investigation, not a causal explanation. Include case counts, uncertainty and other sites. A new outcome coding rule or a changed intervention can mimic model decay. The group audit asks which population carries the change.
Check prediction and action feedback
If the alert causes staff to rescue a shipment, the observed miss rate among alerted cases is partly an effect of the policy. A naive comparison of alerted and unalerted outcomes cannot recover what would have happened without the alert. Preserve action logs and, if a policy experiment is justified, design it through governance rather than treating an observational chart as proof. Causal inference supplies the separate question.
Choose a bounded response
Investigate instrumentation, target definition and route operations before retraining. A fixed model can be perfectly implemented yet wrong for the new process. Test candidate changes with rolling-origin evaluation and an untouched later period; keep the former model available for rollback. The backtest guide describes the clock.
Implementation
# Period, backlog band, original issued alert, mature missed-handoff outcome.
audited_cases = [
("earlier", "middle", 0, 0), ("earlier", "middle", 1, 1),
("earlier", "middle", 0, 0), ("earlier", "middle", 1, 1),
("later", "middle", 0, 1), ("later", "middle", 1, 1),
("later", "middle", 0, 1), ("later", "middle", 1, 1),
]
def period_audit(rows, band):
grouped = {}
for period, observed_band, alert, outcome in rows:
if observed_band != band:
continue
counts = grouped.setdefault(period, {"support": 0, "outcomes": 0, "errors": 0})
counts["support"] += 1
counts["outcomes"] += outcome
counts["errors"] += alert != outcome
return {period: {**counts,
"event_rate": counts["outcomes"] / counts["support"],
"alert_error": counts["errors"] / counts["support"]}
for period, counts in grouped.items()}
report = period_audit(audited_cases, "middle")
assert report["earlier"]["event_rate"] == 0.5
assert report["later"]["event_rate"] == 1.0
assert report["later"]["alert_error"] == 0.5Performance and operating cost
A period-by-slice audit over N mature cases takes O(N) time and O(P times S) summary memory for P periods and S slices. Outcome delay, label audits and policy feedback complicate interpretation far more than computing the rates.
Common Mistakes
- Do not call unresolved outcomes negative.
- Do not treat a within-slice change as proof of a specific cause.
- Do not retrain directly on outcomes altered by the existing intervention without assessing feedback.
