Assign merchants consistently, log the model that served each receipt and enforce operational guardrails before reading delayed outcomes.
Project: run a receipt-model experiment with exposure evidence
Freeze the experiment plan
The current receipt model and candidate r9 share one endpoint. Assign eligible merchants to arms using a stable experiment revision, record the candidate traffic share and specify timeout, unsafe-decision and manual-review stop limits. Keep shadow traffic outside the experimental outcome cohort. Assignment and exposure are separate records because a routed request may use a fallback instead of its assigned model.
Build failure cases
Send repeated receipts from one merchant through different workers, force a candidate timeout, simulate a model digest mismatch and add a merchant that appears in both arms due to an old assignment cache. Record every deviation. Create a shared feature-store slowdown that affects both arms, demonstrating why a per-arm metric alone cannot prove the candidate caused it. Freeze expected stop or hold states before looking at outcome labels.
Monitor and join outcomes
Compare assigned and exposed counts, crossovers, fallback, timeout and manual-review rates by arm. Stop promptly on predefined harm. For the slower quality result, join only mature outcomes by prediction ID and report coverage in both arms. Use the assignment cohort as the primary denominator; inspect actually served digests as a diagnostic. Guardrails and interference keep operational changes from being mistaken for a clean model comparison.
Publish a decision record
The review record names experiment revision, allocation rule, model digests, eligibility, exposure loss, guardrail events, mature outcome window and unresolved limitations. A stopped experiment remains useful evidence but does not establish that the control is universally better. If a candidate is approved, send it through normal promotion and rollback checks rather than directly changing the full-traffic pointer.
Implementation
def exposure_audit(assignments, exposures):
crossover = []
missing = []
for decision_id, assigned in assignments.items():
served = exposures.get(decision_id)
if served is None:
missing.append(decision_id)
elif served not in (assigned, "fallback"):
crossover.append(decision_id)
return {"state": "hold" if crossover or missing else "ready",
"crossover": sorted(crossover), "missing": sorted(missing)}
assigned = {"decision-47": "control", "decision-82": "candidate"}
served = {"decision-47": "control", "decision-82": "fallback"}
assert exposure_audit(assigned, served)["state"] == "ready"
assert exposure_audit(assigned, {**served, "decision-82": "control"})[
"crossover"] == ["decision-82"]
Performance and operating cost
Auditing a assignments costs O(a) expected time and O(a) space in the worst case for flagged IDs. Exposure storage and delayed-label joins scale with traffic. A fallback recorded as an expected exposure state still belongs in the assignment cohort and must be counted in operational and outcome analysis.
Common Mistakes
- Calling a fallback exposure a candidate model prediction.
- Dropping timed-out assigned units from the comparison.
- Ignoring merchants that cross arms between workers.
- Promoting a candidate from an immature or biased outcome slice.
Read next
- Model experiments: separate assignment from actual exposure
- Model experiment guardrails: stop harm without misreading the sample
- Prediction-outcome joins: evaluate only mature, matched decisions
- Promotion evidence: bind evaluation, contract and rollback to one digest
- Shadow and canary rollout: compare a candidate without losing a rollback
Continue the workflow: Project: release a receipt threshold with review evidence.
