A model A/B comparison needs operational stop rules, stable populations and a plan for shared capacity or cross-arm effects.
Model experiment guardrails: stop harm without misreading the sample
Set stop rules before traffic moves
Choose a maximum timeout rate, unsafe-decision count, manual-review load and urgent-case failure boundary before seeing candidate results. Separate a fast safety stop from a slower quality conclusion that needs mature outcomes. State the observation window and minimum count for each rule. Operational alerts identify immediate customer symptoms; an experiment stop also records which arm, model digest and allocation revision produced the breach.
Watch shared systems
Two arms may compete for the same feature store, GPU queue or manual-review team. A candidate that consumes more capacity can slow the control arm, so the control is no longer an untouched baseline. Track total service load, arm-specific load and downstream saturation. If merchants interact or share limits, assign at a cluster that contains the interference boundary where practical. A/B routing does not remove confounding from concurrent infrastructure changes; annotate releases and incidents in the experiment timeline.
Keep quality cohorts comparable
A receipt outcome may mature weeks after exposure. Evaluate both arms on the same eligibility, assignment and outcome-maturity rules. Report missing-label coverage by arm and use intent-to-treat as the primary comparison when fallback or crossover occurs. An apparent candidate win from only reviewed high-risk receipts may reflect selection into review rather than improved model decisions. Mature outcome joins and label revisions keep the evaluation reproducible.
Define a decision that can be defended
Before launch, write the primary outcome, guardrails, sample-size reasoning and stopping policy. After a stop, retain the allocation and exposure logs so a reviewer can separate harmful model output from serving failure. Do not repeatedly peek at an ordinary fixed-sample significance threshold and call the first favorable result final. The project exercises a candidate that raises manual-review load despite an appealing preliminary quality score.
Implementation
def arm_guardrail(control, candidate):
for arm in (control, candidate):
if arm["requests"] < 470:
return {"state": "observe", "reason": "small-arm"}
candidate_timeout = candidate["timeouts"] / candidate["requests"]
candidate_review = candidate["manual_review"] / candidate["requests"]
if candidate["unsafe_decisions"] > 0:
return {"state": "stop", "reason": "unsafe-decision"}
if candidate_timeout > 0.025 or candidate_review > 0.08:
return {"state": "stop", "reason": "operational-guardrail"}
return {"state": "continue"}
control = {"requests": 500, "timeouts": 4,
"manual_review": 23, "unsafe_decisions": 0}
candidate = {"requests": 500, "timeouts": 7,
"manual_review": 47, "unsafe_decisions": 0}
assert arm_guardrail(control, candidate)["state"] == "stop"
Performance and operating cost
This two-arm check is O(1) time and space after aggregation. Stable experiments may need substantial traffic and long follow-up for delayed outcomes; a fast operational stop should not be confused with statistical evidence that one model is better. Shared-capacity effects require system-level metrics in addition to per-arm counters.
Common Mistakes
- Choosing guardrails only after seeing candidate traffic.
- Calling an immature quality estimate a final win.
- Ignoring candidate load that degrades the control arm.
- Dropping fallback decisions from the assigned cohort.
Read next
- Model experiments: separate assignment from actual exposure
- Project: run a receipt-model experiment with exposure evidence
- Prediction-outcome joins: evaluate only mature, matched decisions
- Model alerts: page on customer symptoms with a named owner
- Project: operate receipt scoring with a deadline and overload path
Continue the workflow: Feedback policy shift: compare models when labels depend on routing.
Continue the workflow: Bandit decision logs: preserve action probabilities and delayed rewards.
