Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: decide whether a branch checkout experiment has enough power

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Translate a useful payment lift into branches per arm, then reject an underpowered launch before collecting outcomes.

Freeze the practical question

The checkout team wants a five-percentage-point payment-success improvement from a 20% planning baseline while fraud-review load stays within capacity. Branches receive the interface assignment because staff training is shared. Write the primary outcome, guardrail, analysis population and fixed decision rule before recruiting branches. The useful effect is a cost judgment, not a number selected from a pilot p-value. The individual-unit approximation supplies only the first budget.

Convert to assignment units

The independent-user approximation yields 1,094 observations per arm under the stated planning multipliers. With 20 users per branch, estimated within-branch correlation of 0.03 and 90% outcome retention, the simple inflation suggests 96 branches per arm. Only 58 per arm are available. Show a sensitivity range for correlation and branch size; do not fill the gap by counting multiple sessions per user. The branch adjustment is still approximate but exposes the feasibility problem.

Reject a weak claim, not the investigation

At 58 branches per arm, a negative result would leave the decision-worthy effect unresolved under the planned design. Options include a longer recruitment window, a larger effect threshold justified by economics, or a different valid assignment scheme. Do not launch and reinterpret a nonsignificant result as proof of no benefit. Keep an explicit fraud-review guardrail and avoid adding new primary outcomes later. The stopping rule remains fixed.

Handoff the experiment packet

Record planning baseline, useful effect, allocation, alpha, target power, branch sizes, correlation range, retention, expected outcome maturity and branch availability. Name the decision owner and conditions for a revised design. If the study proceeds later, analyze by assigned branch and report effect sizes and uncertainty. Effect reporting and cluster sampling govern the analysis, not a session-level shortcut.

Implementation

python
from math import ceil

def checkout_plan_gate(independent_users, users_per_branch, correlation,
                       retention, available_branches):
    if independent_users < 1 or users_per_branch < 2 or not 0 <= correlation < 1:
        raise ValueError("invalid study design")
    if not 0 < retention <= 1 or available_branches < 0:
        raise ValueError("invalid retention or branch capacity")
    required = ceil(ceil(independent_users *
                         (1 + (users_per_branch - 1) * correlation) /
                         retention) / users_per_branch)
    return {"required_per_arm": required,
            "state": "ready" if available_branches >= required else "hold:branch-capacity"}

plan = checkout_plan_gate(1094, 20, 0.03, 0.90, 58)
assert plan == {"required_per_arm": 96, "state": "hold:branch-capacity"}

Performance and operating cost

The gate uses O(1) time and space. Recruiting branches, preserving assignment and waiting for mature payment outcomes dominate study cost. An underpowered launch can be costly twice: it consumes users and still fails to distinguish a useful effect from noise.

Common Mistakes

  • Calling 58 branches sufficient because their customer rows exceed the individual target.
  • Moving the useful-effect threshold after a null result.
  • Treating nonsignificance as proof that the screen cannot help.
  • Ignoring fraud-review capacity when planning payment success.

Read next

Continue the workflow: Project: monitor a checkout experiment with declared interim looks.

ai-data
applied-statistics
Storage details