A test asks whether data conflict with a specified null process; the observed difference and its uncertainty answer the practical question.
Hypothesis tests: pair the decision rule with an effect size
Specify the question first
For a proposed review workflow, define the outcome, unit, direction and smallest useful change. State the null process and significance level before results arrive. A p-value is computed under that null; it is not the probability the null is true. A small p-value can accompany a trivial effect in a large dataset, while a useful effect may remain uncertain in a small one.
Use assignment structure
A permutation test swaps labels only where exchangeability is justified. If stores were randomized, shuffle store assignments rather than individual receipts. If the same customer was observed before and after, use a paired statistic or a design-specific model. Randomization units] determine valid comparison mechanics.
Report magnitude and precision
Show the treated and control rates, absolute difference, interval and counts. A 1.2 percentage-point reduction may matter differently at 400 versus four million monthly receipts. Intervals] show uncertainty in the magnitude; a binary significance label cannot replace them.
Exercise a null fixture
Create two groups with the same outcomes but different row order and confirm the statistic is zero. Then inject a known change and check its sign. Verify that a test cannot run on empty groups, missing outcomes or a post-hoc slice chosen because it looked favorable.
Implementation
from itertools import chain
def rate_difference(treatment, control):
if len(treatment) == 0 or len(control) == 0:
raise ValueError("both groups require observed outcomes")
if any(value not in (0, 1) for value in chain(treatment, control)):
raise ValueError("outcomes must be binary")
return sum(treatment) / len(treatment) - sum(control) / len(control)
observed_difference = rate_difference(treatment_outcomes, control_outcomes)Performance and operating cost
Calculating a rate difference is O(N). A permutation test with B resamples costs O(BN) in a direct implementation. The inferential validity depends more on assignment and missingness than on the speed of arithmetic.
Common Mistakes
- Do not call a p-value the probability the null is true.
- Do not omit the effect size and group denominators.
- Do not shuffle receipt rows when assignment happened by store.
Read next
- Experiment design: assign the right unit and guard against interference
- Multiple comparisons and peeking: protect a predeclared decision rule
- Confidence intervals: interpret coverage and precision honestly
- Metric denominators and cohorts: make a rate reproducible
Continue the workflow: Paired randomization checks: enumerate sign assignments under a sharp null.
Continue the workflow: Prior sensitivity: show when limited data leave a decision exposed.
Continue the workflow: Minimum detectable effect: plan a study around a useful change.
Continue the workflow: Equivalence margins: require the whole interval to fit.
