Equivalence testing asks whether a difference is small enough in both directions under limits chosen before analysis.
Equivalence margins: require the whole interval to fit
Define the quantity and both limits
A revised invoice-review screen should keep average review duration close enough to the incumbent that either direction of change is operationally acceptable. Define candidate minus incumbent in seconds, specify the eligible invoice population and commit to lower and upper acceptable differences before reading the pilot. A symmetric margin is convenient, but asymmetric limits can express different costs for a faster and a slower screen. The estimand and sampling frame fix what the difference represents.
Reverse the usual burden of proof
A conventional difference test begins with zero difference as the null. Failing to reject it says only that the data are inconclusive for that test. An equivalence claim instead needs evidence that the true difference exceeds the lower limit and stays below the upper limit. At a five-percent level, a compatible two-one-sided procedure can be read from a valid two-sided ninety-percent confidence interval: the entire interval must fit strictly between both predeclared limits. The function below checks containment; it does not calculate the interval. Coverage assumptions still matter.
Match uncertainty to assignment
If the same reviewer handles both interfaces on matched invoice batches, the interval must use within-reviewer differences rather than independent-row variance. If the pilot assigns whole branches, branch-level dependence belongs in the design and interval. Attrition, crossovers and a changed review clock can pull groups toward apparent similarity. Report both the assigned-group analysis and an adherence-sensitive analysis when those disruptions occur, with reasons for exclusions. Pairing and cluster resampling offer distinct routes, depending on assignment.
Deliver a bounded claim
Successful containment supports equivalence only for the named endpoint, margin, population and observation window. It does not establish identical workflows, equal worst-case times or equivalence for error rates. Show the estimated difference, interval, limits, sample counts and missingness together. A wide interval crossing a limit is inconclusive even when its midpoint is near zero. A one-sided harm limit answers another release question; the project joins the checks.
Implementation
def equivalence_from_interval(lower_ci, upper_ci, lower_margin, upper_margin):
if lower_ci > upper_ci or lower_margin >= upper_margin:
raise ValueError("ordered interval and margins required")
return lower_margin < lower_ci and upper_ci < upper_margin
assert equivalence_from_interval(-1.4, 1.8, -2.5, 2.5)
assert not equivalence_from_interval(-2.7, 1.2, -2.5, 2.5)
Performance and operating cost
Interval containment is O(1) time and space. Constructing a valid interval can require O(n) data passes or repeated cluster resampling, while reviewer-level matching and incomplete cases must be audited first. The four comparisons cannot repair a biased pilot or a margin chosen after seeing results.
Common Mistakes
- Calling a nonsignificant difference test proof of equivalence.
- Changing the margin after seeing the estimate.
- Using an invalid independent-row interval for a branch-assigned pilot.
- Claiming every workflow outcome is equivalent from one average-duration endpoint.
Read next
- Noninferiority: orient the harm margin before examining an interval
- Project: decide whether an invoice workflow preserves speed and safety
- Confidence intervals: interpret coverage and precision honestly
- Paired comparisons: analyze within-unit changes and preserve the match
- Minimum detectable effect: plan a study around a useful change
Continue the workflow: Project: decide whether a handheld parcel scale can replace the dock scale.
