A live model comparison needs stable assignment, a record of which model actually served, and a joinable outcome clock.
Model experiments: separate assignment from actual exposure
Choose the randomization unit
A receipt-risk experiment may assign by merchant rather than individual receipt so repeated decisions for one merchant stay on the same model. Hash a stable pseudonymous unit ID with an experiment revision and salt, then keep the allocation fixed for that revision. Assigning by request ID can send related receipts to both arms and contaminate the comparison. Record eligibility rules before assignment. Shadow traffic is different: its candidate output is hidden and cannot measure the customer effect of a model decision.
Log actual exposure after routing
Assignment means a unit was allocated; exposure means the selected model actually served an eligible decision. A timeout, fallback, cache hit or previous model retained by a long-running worker can prevent the intended exposure. Store experiment ID, unit ID, assigned arm, served model digest, decision ID and timestamp. Avoid logging raw receipt payloads. Inference logging provides a restricted join key, while outcome joins later connect exposed decisions to mature labels.
Detect crossovers and denominator drift
Monitor eligible units, assigned units, exposed units, crossovers and missing exposure records by arm. If one model times out more often, analyzing only successful exposures creates a selected population. Keep the intent-to-treat assignment cohort for the primary experiment estimate and report served-model diagnostics separately. Preserve a mapping from unit to arm for the experiment duration. If allocation changes midstream, use a new experiment revision instead of rewriting the old assignment history.
Test stable behavior
Replay the same merchant across workers and days, change request IDs, and verify arm stability. Force a timeout to confirm assignment remains logged while actual exposure records fallback. The receipt experiment project uses those records to detect crossovers and compare operational guardrails before it considers delayed quality outcomes.
Implementation
from hashlib import sha256
def assign_model(experiment_id, merchant_key, candidate_share=0.23):
if not 0 <= candidate_share <= 1:
raise ValueError("candidate share must be a fraction")
identity = (experiment_id + ":" + merchant_key).encode()
bucket = int.from_bytes(sha256(identity).digest()[:8], "big") / 2**64
return "candidate" if bucket < candidate_share else "control"
first = assign_model("receipt-risk-r9", "merchant-47")
assert first == assign_model("receipt-risk-r9", "merchant-47")
assert assign_model("receipt-risk-r9", "merchant-47", 0) == "control"
assert assign_model("receipt-risk-r9", "merchant-47", 1) == "candidate"
Performance and operating cost
Assignment hashes k identity bytes in O(k) time with O(1) digest space. Persisted exposure and fallback events add storage proportional to eligible decisions. Stable hashing is not itself an experiment analysis: eligibility, missing exposures and outcome delay still determine whether the comparison is interpretable.
Common Mistakes
- Randomizing each receipt when merchant-level spillover is expected.
- Equating assigned arm with model actually served.
- Analyzing only successful exposures after one arm times out more often.
- Changing allocation without a new experiment revision.
Read next
- Model experiment guardrails: stop harm without misreading the sample
- Project: run a receipt-model experiment with exposure evidence
- Shadow and canary rollout: compare a candidate without losing a rollback
- Inference logs: keep diagnostic joins without copying sensitive payloads
- Prediction-outcome joins: evaluate only mature, matched decisions
