Design a review-workflow experiment with a fixed population, assignment unit, primary rate, uncertainty method and stopping rule.
Project: evaluate a receipt-review workflow without changing the question midstream
Write the analysis contract
Define eligible receipts, assignment unit, outcome window, primary metric and smallest useful improvement before simulating treatment. Include an outage day, a repeated customer and delayed outcomes in the fixture. State whether analysis compares assigned or exposed units. Population definition] is the first deliverable.
Check randomization and data
Use a stable hash assignment, then verify no unit appears in both variants. Reconcile raw, excluded, pending and measured outcomes. Report group counts and pre-treatment balance. If assignment happened by customer, calculate uncertainty by customer rather than pretending each receipt is independent. Cluster resampling] supplies the uncertainty exercise.
Apply a frozen rule
Estimate the difference in review rates and its interval. Compare the result with the declared improvement target and guardrails. Show a predeclared end condition and label exploratory slices separately. Multiplicity] must be addressed before inspecting a collection of favorable segments.
Submit a decision record
Provide fixture, analysis code, manifest, outcome counts, interval, assumption checks and a one-page decision. Include a failure drill: duplicate assignment, missing store partition or immature label must block a clean result. A chart without numerator, denominator and exclusion counts does not complete the project.
Implementation
def experiment_integrity(assignments, outcomes):
variant_by_customer = {}
for customer_id, variant in assignments:
previous = variant_by_customer.setdefault(customer_id, variant)
if previous != variant:
raise ValueError("customer assigned to both variants")
known_customers = set(variant_by_customer)
if any(customer_id not in known_customers for customer_id in outcomes):
raise ValueError("outcome without assignment")
return len(known_customers)Performance and operating cost
Assignment integrity is O(N) time and O(K) memory for N events and K customers. Cluster bootstrap adds repeated computation; documenting excluded and pending outcomes adds modest storage.
Common Mistakes
- Do not change the primary metric after viewing the result.
- Do not treat pending outcomes as failures.
- Do not report receipt-level precision for customer-level assignment.
