A challenger can score the same incoming cases as the live model without changing decisions, allowing paired comparison when outcomes later mature.
Champion–challenger shadow comparison
Keep the live action fixed
A challenger predicts missed handoffs beside the current champion, but only the champion’s alert drives dispatch. Log both scores, versions, input snapshot, champion action and later mature outcome under one case ID. This avoids quietly changing a production policy during measurement. The shadow outcome still reflects the champion’s intervention, so it may not equal the outcome the challenger policy would have caused. Policy feedback explains the limitation.
Pair each comparison
Compare loss differences case by case; the code reports the mean challenger-minus-champion Brier loss on the same four cases. This removes one source of variation from comparing two unrelated cohorts. It is not a proof of better future policy value, and four cases are too few for a release. Paired resampling can quantify uncertainty with the right independent unit.
Audit latency and missingness
The challenger may fail to score hard cases, score after the action deadline or require a feature unavailable to the live model. Include all attempted cases and report missing-score rate and latency, not only rows where both outputs exist. For a fair candidate comparison, preserve the same decision-time feature contract. Feature timing is non-negotiable.
Translate score gain into an action policy
A lower proper scoring loss does not automatically improve dispatch. Compare predeclared thresholds, false-negative cost, calibration, subgroup error and review load on a held-out future period. If challenger alerts would change outcomes, an approved prospective test may be needed to estimate policy impact. Threshold costing and group auditing state the minimum checks.
Name promotion and rollback rules
Record the minimum support, acceptable regression bounds, missingness ceiling, operational owner and rollback trigger before viewing the shadow result. A challenger can remain in shadow until enough outcomes mature. The project closes with a bounded pilot decision, not an automatic promotion.
Implementation
# Same case, two issued risk estimates, then the mature outcome.
shadow_cases = [
("S741", 0.18, 0.12, 0), ("S742", 0.62, 0.79, 1),
("S743", 0.54, 0.37, 0), ("S744", 0.71, 0.84, 1),
]
def paired_brier_delta(rows):
differences = []
for shipment_id, champion_risk, challenger_risk, outcome in rows:
if not (0 <= champion_risk <= 1 and 0 <= challenger_risk <= 1):
raise ValueError("risk outside probability range")
champion_loss = (champion_risk - outcome) ** 2
challenger_loss = (challenger_risk - outcome) ** 2
differences.append(challenger_loss - champion_loss)
return sum(differences) / len(differences)
delta = paired_brier_delta(shadow_cases)
assert delta < 0
assert len({shipment_id for shipment_id, *_ in shadow_cases}) == 4Performance and operating cost
Paired scoring of N logged cases costs O(N) time and O(1) running memory; the teaching code stores O(N) differences. Shadow inference adds model serving cost and latency without reducing live-model cost. Mature outcomes and a policy-effect evaluation may take much longer than scoring.
Common Mistakes
- Do not let a shadow score silently change the champion action.
- Do not evaluate only cases with successful challenger scores.
- Do not equate improved scoring loss under the champion policy with proven challenger policy value.
Read next
- Rolling-origin retraining with an outcome embargo
- Paired bootstrap intervals for model gain
- Decision thresholds: choose an action from probabilities and error costs
- Distribution shift response project
Continue the workflow: Compression Pareto review and shadow check.
