A sample-size plan joins a decision-worthy effect, error rates, baseline variability and the actual assignment unit.
Minimum detectable effect: plan a study around a useful change
Start with the smallest useful change
A checkout team can justify shipping a new payment screen if successful payment rises from 20% to at least 25% without increasing fraud review. That five-percentage-point difference is a practical threshold chosen from business cost, not from whichever result later becomes significant. State the target population, primary binary outcome, equal or unequal allocation, two-sided false-positive rate and desired power before exposure begins. Effect size is part of the decision contract.
Use an approximation transparently
For two independent equally sized groups, a normal approximation can estimate required users per arm from a planning baseline, target difference and critical multipliers. The code uses a pooled planning variance and rounds up. It is a starting budget, not an exact guarantee of achieved power: actual baseline rates, dropout, repeated users and a discrete test can change the needed size. Recalculate with design-specific software before committing traffic. Assignment units may be accounts rather than sessions.
Plan a range, not one magic number
Show sample needs for several plausible baseline rates and effect thresholds. Smaller effects cost sharply more observations because the difference enters the denominator squared in this approximation. If capacity cannot reach the useful effect, change the question, extend the run, or acknowledge that the study will be imprecise. Do not quietly redefine the effect after a null result. Cluster and attrition adjustments translate a user-level target into a deployable plan.
Protect the resulting decision rule
Freeze success criteria, launch window, inclusion rules and any interim looks. Adding many outcomes or repeated unplanned significance checks changes the false-positive behavior. Power is the chance of rejecting a specified false null under an assumed data-generating scenario; it is not the chance that the eventual hypothesis is true. Stopping rules and the project carry the plan into execution.
Implementation
from math import ceil
def approximate_users_per_arm(baseline_rate, target_rate,
critical_alpha=1.96, critical_power=0.84):
if not 0 < baseline_rate < 1 or not 0 < target_rate < 1:
raise ValueError("rates must be between zero and one")
difference = abs(target_rate - baseline_rate)
if difference == 0:
raise ValueError("target must differ from baseline")
pooled_rate = (baseline_rate + target_rate) / 2
variance = pooled_rate * (1 - pooled_rate)
return ceil(2 * (critical_alpha + critical_power) ** 2 *
variance / difference ** 2)
needed = approximate_users_per_arm(0.20, 0.25)
assert needed == 1094
assert approximate_users_per_arm(0.20, 0.23) > needed
Performance and operating cost
The approximation is O(1) time and space. Actual study cost scales with independent assigned units, outcome maturity and review capacity. A power number computed from session rows instead of the randomized account unit can be far too optimistic despite correct arithmetic.
Common Mistakes
- Choosing the minimum effect after inspecting the result.
- Treating approximate sample size as a guarantee of achieved power.
- Counting sessions when assignment and dependence are at account level.
- Ignoring the other primary outcome or repeated interim tests in the design.
Read next
- Cluster and attrition planning: count independent branches, not rows
- Project: decide whether a branch checkout experiment has enough power
- Hypothesis tests: pair the decision rule with an effect size
- Experiment design: assign the right unit and guard against interference
- Multiple comparisons and peeking: protect a predeclared decision rule
Continue the workflow: Noninferiority: orient the harm margin before examining an interval.
Continue the workflow: Early stopping: separate the success decision from effect estimation.
