An experiment compares outcomes for stable groups under a defined eligibility rule. Assignment is the group a unit receives; exposure is the point at which that unit can actually encounter the treatment. Logging assignment at page load when the feature is below the fold can count people who never saw it. Logging exposure only after a variant-specific action can bias the groups. The analysis unit, enrollment rule, primary outcome, observation window, and guardrail metrics should be fixed before launch. A feature flag can deliver variants, but a rollout percentage alone does not produce a sound experiment.
Experiment Assignment, Exposure, and Metric Integrity
Working case
The case-review team tests whether the new assignment panel reduces time to assign a case. Reviewers are assigned by stable account ID, but only those who open an editable case enter the eligible population. Exposure is recorded when either panel is rendered, using the same rule in control and treatment. The result event contains an experiment ID, variant, and coarse completion time but no case title. When the treatment panel crashes before exposure logging, observed groups become unbalanced. The team pauses interpretation, investigates the missing events, and keeps error rate as a guardrail rather than celebrating a faster average among survivors.
Implementation boundary
function shouldRecordExposure(eligible, rendered, alreadyRecorded) {
return eligible && rendered && !alreadyRecorded;
}
console.log(shouldRecordExposure(true, false, false));
// Output: falseUse a stable unit that matches the product decision: account, organization, or session, with a documented rule for shared devices and multiple accounts. Keep eligibility identical across variants and record one exposure per unit per experiment context when the experience is actually presented. Deduplicate retries and carry the same unit ID into outcome events. Define one primary metric and a small set of guardrails before observing results. Check traffic allocation, missing events, crossovers, and sample-ratio mismatch; persistent imbalance can signal assignment or logging defects. Avoid changing variants mid-experiment without recording a new analysis period. Minimize telemetry fields and honor optional-measurement policy where it applies.
Cost and boundaries
Stable assignment can be O(1) hashing or a stored lookup, while event collection scales with exposures and outcomes. Dedupe state adds storage proportional to the enrolled units. Measuring at a different point for each variant creates bias that no larger sample fixes. Very small groups produce noisy estimates; deciding a winner after each refresh increases false-positive risk. Analysis also costs engineering time to inspect guardrails and data quality. Track exposure-to-outcome join rate, missing identifiers, sample ratio, crossover, and error rate before examining the primary result.
Failure trace
The treatment records exposure when its button is clicked, while control records it when the panel appears. People who never click disappear from the treatment denominator, making the variant seem faster. Move both exposure points to equivalent presentation. Another client assigns by a random number on every reload; a reviewer sees both panels and cannot be analyzed as one unit. Test account switching, route reload, offline event replay, duplicate exposure, variant-specific crash, cohort imbalance, and a privacy setting that disables optional measurements.
Verification
- Control and treatment use one equivalent exposure point.
- Outcomes join to a stable assigned unit.
- Imbalance and crossover are checked before interpreting results.
Practice drill
Predeclare an eligible reviewer population, unit ID, exposure point, assignment split, primary completion metric, and error guardrail. Simulate 47 eligible reviewers, with some leaving before panel render and one treatment crash. Verify only rendered panels count as exposed under both variants. Duplicate two events and confirm dedupe. Change the assignment seed in a test and detect crossover. Inspect exposure counts for unexpected imbalance before reading any claimed improvement.
Decision note
Measure comparable exposed groups under a stable assignment and verify data quality before interpreting an outcome.
Common Mistakes
- Calling every page load an exposure.
- Changing the assignment unit during a run.
- Treating a rollout percentage as a complete experiment.
Connected lessons
Feature Release and Experiment Controls; Feature Flag Evaluation Scope and Safe Default; Gradual Rollout, Kill Switch, and State Compatibility; Flag Telemetry, Audit, and Retirement; Telemetry Minimization and Retention; User Journey SLOs and Burn Alerts; Quality and Capacity Engineering.
Apply and check
Build Project: assignment panel release experiment and review Web Development: release and GraphQL decisions quiz.
