A one-sided tolerance can rule out an unacceptable loss if the endpoint direction and analyzed population remain fixed.
Noninferiority: orient the harm margin before examining an interval
Put harm on the positive side
A new review interface may save staffing cost while allowing at most a small increase in missed duplicate invoices. Define the endpoint as candidate miss rate minus incumbent miss rate, so positive differences are harmful. A margin of 0.004 means four additional misses per thousand eligible invoices are the largest acceptable loss; that number must come from the operating decision, not the observed result. The unit and adjudication window must be identical across arms. Miss-rate definitions require audited unflagged invoices.
Test the boundary at issue
For this orientation, noninferiority requires a valid upper confidence bound for the difference below the positive harm margin. A point estimate below the margin is insufficient when its uncertainty crosses the boundary. If higher values are beneficial, reverse the contrast and use the corresponding lower bound instead. State whether the bound is one-sided and what error level it represents; do not silently substitute a two-sided interval with a different confidence level. Equivalence controls both directions, not only harm.
Audit who remained in the comparison
An assigned-group analysis answers the effect of offering the new interface under pilot assignment, while a protocol-adherent analysis describes use among people who followed the workflow. Crossovers and exclusions can make groups look similar, which is dangerous when similarity is the desired finding. Keep an assignment ledger, predefine material deviations, show both analyses and explain disagreements instead of selecting the easier result. If review outcomes are still pending, delay the decision. The assignment unit determines the uncertainty structure.
Separate tradeoffs from guardrails
A smaller staff cost does not compensate automatically for a miss-rate interval that still permits unacceptable harm. Report the cost saving and its uncertainty separately from the safety boundary. Add a high-severity error check and missing-adjudication table, because a single average miss rate may hide concentrated damage. The outcome is scoped to the pilot frame and period. The invoice project checks duration equivalence and error noninferiority together; the decision-rule lesson handles added endpoints.
Implementation
def noninferior_on_harm_scale(upper_difference_bound, harm_margin):
if harm_margin <= 0:
raise ValueError("harm margin must be positive")
return upper_difference_bound < harm_margin
# Difference is candidate minus incumbent missed-duplicate rate.
assert noninferior_on_harm_scale(0.0026, 0.004)
assert not noninferior_on_harm_scale(0.0051, 0.004)
Performance and operating cost
The boundary check is O(1). Obtaining a defensible bound costs at least O(n) data work and may require branch-level variance estimation. Counting only promptly adjudicated cases is cheaper but can bias the contrast toward noninferiority when difficult cases mature later.
Common Mistakes
- Changing the sign of a difference without changing the decision rule.
- Comparing only the point estimate with the margin.
- Dropping crossovers or pending cases without a declared analysis plan.
- Claiming no harm when the bound allows a loss beyond the agreed tolerance.
Read next
- Equivalence margins: require the whole interval to fit
- Project: decide whether an invoice workflow preserves speed and safety
- Experiment design: assign the right unit and guard against interference
- Multiple comparisons and peeking: protect a predeclared decision rule
- Diagnostic performance: separate sensitivity, specificity and predictive value
