Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Minimum detectable effect: plan a study around a useful change

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A sample-size plan joins a decision-worthy effect, error rates, baseline variability and the actual assignment unit.

Start with the smallest useful change

A checkout team can justify shipping a new payment screen if successful payment rises from 20% to at least 25% without increasing fraud review. That five-percentage-point difference is a practical threshold chosen from business cost, not from whichever result later becomes significant. State the target population, primary binary outcome, equal or unequal allocation, two-sided false-positive rate and desired power before exposure begins. Effect size is part of the decision contract.

Use an approximation transparently

For two independent equally sized groups, a normal approximation can estimate required users per arm from a planning baseline, target difference and critical multipliers. The code uses a pooled planning variance and rounds up. It is a starting budget, not an exact guarantee of achieved power: actual baseline rates, dropout, repeated users and a discrete test can change the needed size. Recalculate with design-specific software before committing traffic. Assignment units may be accounts rather than sessions.

Plan a range, not one magic number

Show sample needs for several plausible baseline rates and effect thresholds. Smaller effects cost sharply more observations because the difference enters the denominator squared in this approximation. If capacity cannot reach the useful effect, change the question, extend the run, or acknowledge that the study will be imprecise. Do not quietly redefine the effect after a null result. Cluster and attrition adjustments translate a user-level target into a deployable plan.

Protect the resulting decision rule

Freeze success criteria, launch window, inclusion rules and any interim looks. Adding many outcomes or repeated unplanned significance checks changes the false-positive behavior. Power is the chance of rejecting a specified false null under an assumed data-generating scenario; it is not the chance that the eventual hypothesis is true. Stopping rules and the project carry the plan into execution.

Implementation

python
from math import ceil

def approximate_users_per_arm(baseline_rate, target_rate,
                              critical_alpha=1.96, critical_power=0.84):
    if not 0 < baseline_rate < 1 or not 0 < target_rate < 1:
        raise ValueError("rates must be between zero and one")
    difference = abs(target_rate - baseline_rate)
    if difference == 0:
        raise ValueError("target must differ from baseline")
    pooled_rate = (baseline_rate + target_rate) / 2
    variance = pooled_rate * (1 - pooled_rate)
    return ceil(2 * (critical_alpha + critical_power) ** 2 *
                variance / difference ** 2)

needed = approximate_users_per_arm(0.20, 0.25)
assert needed == 1094
assert approximate_users_per_arm(0.20, 0.23) > needed

Performance and operating cost

The approximation is O(1) time and space. Actual study cost scales with independent assigned units, outcome maturity and review capacity. A power number computed from session rows instead of the randomized account unit can be far too optimistic despite correct arithmetic.

Common Mistakes

  • Choosing the minimum effect after inspecting the result.
  • Treating approximate sample size as a guarantee of achieved power.
  • Counting sessions when assignment and dependence are at account level.
  • Ignoring the other primary outcome or repeated interim tests in the design.

Read next

Continue the workflow: Noninferiority: orient the harm margin before examining an interval.

Continue the workflow: Early stopping: separate the success decision from effect estimation.

ai-data
applied-statistics
Storage details