Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Model experiments: separate assignment from actual exposure

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A live model comparison needs stable assignment, a record of which model actually served, and a joinable outcome clock.

Choose the randomization unit

A receipt-risk experiment may assign by merchant rather than individual receipt so repeated decisions for one merchant stay on the same model. Hash a stable pseudonymous unit ID with an experiment revision and salt, then keep the allocation fixed for that revision. Assigning by request ID can send related receipts to both arms and contaminate the comparison. Record eligibility rules before assignment. Shadow traffic is different: its candidate output is hidden and cannot measure the customer effect of a model decision.

Log actual exposure after routing

Assignment means a unit was allocated; exposure means the selected model actually served an eligible decision. A timeout, fallback, cache hit or previous model retained by a long-running worker can prevent the intended exposure. Store experiment ID, unit ID, assigned arm, served model digest, decision ID and timestamp. Avoid logging raw receipt payloads. Inference logging provides a restricted join key, while outcome joins later connect exposed decisions to mature labels.

Detect crossovers and denominator drift

Monitor eligible units, assigned units, exposed units, crossovers and missing exposure records by arm. If one model times out more often, analyzing only successful exposures creates a selected population. Keep the intent-to-treat assignment cohort for the primary experiment estimate and report served-model diagnostics separately. Preserve a mapping from unit to arm for the experiment duration. If allocation changes midstream, use a new experiment revision instead of rewriting the old assignment history.

Test stable behavior

Replay the same merchant across workers and days, change request IDs, and verify arm stability. Force a timeout to confirm assignment remains logged while actual exposure records fallback. The receipt experiment project uses those records to detect crossovers and compare operational guardrails before it considers delayed quality outcomes.

Implementation

python
from hashlib import sha256

def assign_model(experiment_id, merchant_key, candidate_share=0.23):
    if not 0 <= candidate_share <= 1:
        raise ValueError("candidate share must be a fraction")
    identity = (experiment_id + ":" + merchant_key).encode()
    bucket = int.from_bytes(sha256(identity).digest()[:8], "big") / 2**64
    return "candidate" if bucket < candidate_share else "control"

first = assign_model("receipt-risk-r9", "merchant-47")
assert first == assign_model("receipt-risk-r9", "merchant-47")
assert assign_model("receipt-risk-r9", "merchant-47", 0) == "control"
assert assign_model("receipt-risk-r9", "merchant-47", 1) == "candidate"

Performance and operating cost

Assignment hashes k identity bytes in O(k) time with O(1) digest space. Persisted exposure and fallback events add storage proportional to eligible decisions. Stable hashing is not itself an experiment analysis: eligibility, missing exposures and outcome delay still determine whether the comparison is interpretable.

Common Mistakes

  • Randomizing each receipt when merchant-level spillover is expected.
  • Equating assigned arm with model actually served.
  • Analyzing only successful exposures after one arm times out more often.
  • Changing allocation without a new experiment revision.

Read next

ai-data
mlops
Storage details