An experiment needs a stable assignment unit, eligible population, exposure rule and outcome window before traffic is divided.
Experiment design: assign the right unit and guard against interference
Choose the assignment unit
If reviewers share a queue, assigning individual receipts to two workflows can create interference: one treatment changes wait time for controls. Randomize at reviewer, queue or store level when necessary, then analyze uncertainty at that level. A customer-level assignment may be required when the same customer submits several receipts. Cluster uncertainty] follows the design.
Define exposure
A unit assigned to treatment may never see the new workflow. State whether the primary estimate compares assigned groups or only those exposed. Filtering to exposed users after randomization can reintroduce selection bias. Keep assignment events, exposure events and outcome windows separate. Record when a unit can enter or leave the experiment.
Protect concurrent changes
A marketing campaign, outage or policy update during the test can change receipt mix. Random assignment balances many differences in expectation, but operational incidents still need tracking. Decide whether to pause, extend or analyze incident periods under a written rule. The target population] should match the rollout population.
Run integrity checks
Before reading outcomes, verify stable allocation by unit, no duplicate assignments, balanced pre-treatment characteristics and correct exposure logging. Simulate a customer with two receipts and a store crossing the rollout boundary. Both receipts must receive the intended consistent assignment if customer is the unit.
Implementation
import hashlib
def assign_review_variant(customer_id, experiment_key):
if not customer_id:
raise ValueError("customer identity required")
assignment_key = f"{experiment_key}:{customer_id}".encode("utf-8")
bucket = int.from_bytes(hashlib.sha256(assignment_key).digest()[:4], "big")
return "new-review" if bucket % 2 == 0 else "current-review"Performance and operating cost
Stable hashing is O(length of assignment key) per unit. Cluster assignment reduces the number of independent units and may require a longer experiment than receipt-level assignment.
Common Mistakes
- Do not change assignment per request for the same unit.
- Do not analyze only exposed units without naming the selection problem.
- Do not ignore shared-queue interference.
Read next
- Hypothesis tests: pair the decision rule with an effect size
- Standard error and cluster bootstrap: resample the independent unit
- Multiple comparisons and peeking: protect a predeclared decision rule
- Project: evaluate a receipt-review workflow without changing the question midstream
Continue the workflow: Randomized assignment: check balance and retain every assigned unit.
Continue the workflow: Experiment contract: decision, estimand and assignment unit.
