Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: run a receipt-model experiment with exposure evidence

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Assign merchants consistently, log the model that served each receipt and enforce operational guardrails before reading delayed outcomes.

Freeze the experiment plan

The current receipt model and candidate r9 share one endpoint. Assign eligible merchants to arms using a stable experiment revision, record the candidate traffic share and specify timeout, unsafe-decision and manual-review stop limits. Keep shadow traffic outside the experimental outcome cohort. Assignment and exposure are separate records because a routed request may use a fallback instead of its assigned model.

Build failure cases

Send repeated receipts from one merchant through different workers, force a candidate timeout, simulate a model digest mismatch and add a merchant that appears in both arms due to an old assignment cache. Record every deviation. Create a shared feature-store slowdown that affects both arms, demonstrating why a per-arm metric alone cannot prove the candidate caused it. Freeze expected stop or hold states before looking at outcome labels.

Monitor and join outcomes

Compare assigned and exposed counts, crossovers, fallback, timeout and manual-review rates by arm. Stop promptly on predefined harm. For the slower quality result, join only mature outcomes by prediction ID and report coverage in both arms. Use the assignment cohort as the primary denominator; inspect actually served digests as a diagnostic. Guardrails and interference keep operational changes from being mistaken for a clean model comparison.

Publish a decision record

The review record names experiment revision, allocation rule, model digests, eligibility, exposure loss, guardrail events, mature outcome window and unresolved limitations. A stopped experiment remains useful evidence but does not establish that the control is universally better. If a candidate is approved, send it through normal promotion and rollback checks rather than directly changing the full-traffic pointer.

Implementation

python
def exposure_audit(assignments, exposures):
    crossover = []
    missing = []
    for decision_id, assigned in assignments.items():
        served = exposures.get(decision_id)
        if served is None:
            missing.append(decision_id)
        elif served not in (assigned, "fallback"):
            crossover.append(decision_id)
    return {"state": "hold" if crossover or missing else "ready",
            "crossover": sorted(crossover), "missing": sorted(missing)}

assigned = {"decision-47": "control", "decision-82": "candidate"}
served = {"decision-47": "control", "decision-82": "fallback"}
assert exposure_audit(assigned, served)["state"] == "ready"
assert exposure_audit(assigned, {**served, "decision-82": "control"})[
    "crossover"] == ["decision-82"]

Performance and operating cost

Auditing a assignments costs O(a) expected time and O(a) space in the worst case for flagged IDs. Exposure storage and delayed-label joins scale with traffic. A fallback recorded as an expected exposure state still belongs in the assignment cohort and must be counted in operational and outcome analysis.

Common Mistakes

  • Calling a fallback exposure a candidate model prediction.
  • Dropping timed-out assigned units from the comparison.
  • Ignoring merchants that cross arms between workers.
  • Promoting a candidate from an immature or biased outcome slice.

Read next

Continue the workflow: Project: release a receipt threshold with review evidence.

ai-data
mlops
Storage details