Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Cluster and attrition planning: count independent branches, not rows

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A clustered rollout needs enough independent assignment units after outcome loss, not merely enough observed customer rows.

Identify why rows are dependent

A checkout interface is rolled out by branch because staff training and shared devices make per-customer randomization impractical. Customers within a branch share local procedures, so the effective information is smaller than a same-size independent-user experiment. An intracluster correlation estimate from earlier comparable branches can guide planning, but it has uncertainty. The randomization unit governs both analysis and sample size.

Build a transparent first budget

For roughly equal cluster size m, the simple design-effect approximation is one plus (m minus one) times the intracluster correlation. Multiply the independent-user target by that factor, adjust for expected outcome retention, then round up to whole branches. This does not replace a proper cluster trial calculation: unequal sizes, few branches, covariate adjustment and contamination can change power. The base target must be tied to a useful change before inflation.

Handle missing outcomes explicitly

If payment success is unobserved for some assigned accounts, a retention adjustment budgets more assignment but does not remove bias. Track loss by arm, branch and baseline risk. The analysis should retain the assigned population under its predeclared missing-outcome rule. If one branch is closed mid-study, dropping it can break balance and shrink the true number of independent units. Missingness mechanisms require separate sensitivity work.

Check feasibility and uncertainty

List required branches per arm, eligible accounts per branch, expected retention and the uncertainty around the correlation assumption. Run a range of correlation values; if branch capacity cannot meet the plan, state that the desired effect is not resolvable under this rollout. Avoid a false precision claim from thousands of customers in only a handful of branches. Cluster-aware uncertainty belongs in analysis; the project holds a branch-limited launch.

Implementation

python
from math import ceil

def branches_per_arm(independent_users, users_per_branch, correlation, retention):
    if independent_users < 1 or users_per_branch < 2:
        raise ValueError("positive users and branch size required")
    if not 0 <= correlation < 1 or not 0 < retention <= 1:
        raise ValueError("invalid correlation or retention")
    design_effect = 1 + (users_per_branch - 1) * correlation
    assigned_users = ceil(independent_users * design_effect / retention)
    return ceil(assigned_users / users_per_branch)

branches = branches_per_arm(1094, 20, 0.03, 0.90)
assert branches == 96
assert branches_per_arm(1094, 20, 0.06, 0.90) > branches

Performance and operating cost

The planning calculation is O(1) time and space. Operational cost rises with the number of independent branches, not just customer observations. Larger within-branch correlation and lower retention can make a branch rollout infeasible; collecting more customers in the same few branches does not create new independent assignments.

Common Mistakes

  • Inflating user count while leaving the number of assigned branches unchanged.
  • Treating one pilot estimate of intracluster correlation as exact.
  • Calling attrition adjustment a solution to selective missing outcomes.
  • Applying an individual-level significance test to a branch-randomized release.

Read next

ai-data
applied-statistics
Storage details