A sign check uses the direction of nonzero paired changes and makes its handling of exact ties explicit.
Paired sign checks: treat zeros and direction as design decisions
Keep the match intact
The same support agent handles matched shifts under two ticket views. Define one candidate-minus-incumbent duration per complete agent-shift pair, with negative values favoring the candidate. A zero difference contributes no direction to a conventional sign test, so report how many zero pairs were removed and why they occurred. Rounded time stamps can create artificial ties. Pair construction precedes any test statistic.
Distinguish sign from signed rank
A sign check discards change magnitude and tests whether positive and negative directions are balanced under its null assumptions. A signed-rank procedure orders absolute nonzero differences and uses more magnitude information, but its familiar location interpretation needs a suitable symmetry condition on paired differences. Neither procedure turns a before/after observation into a randomized intervention. Show the differences, not just one test result. Assignment-based sign flips depend on the actual randomization scheme.
Calculate small-sample evidence exactly
For n nonzero pairs under a fair independent-sign null, the number of positive signs has a binomial distribution with probability one half. A two-sided tail can be enumerated with integer combinations. The code reports a two-sided p-value using the smaller directional count; it is deliberately separate from an effect estimate. With very few pairs, attainable p-values are coarse. If repeated shifts share an agent, independent signs can fail even when the pairs themselves are complete. Dependence review remains necessary.
Make the release question practical
Report counts of improvements, deteriorations and ties, plus a meaningful time difference and high-percentile delay. Predeclare whether a zero should count as a neutral outcome or a service success for the operational decision. A small p-value cannot replace a minimum useful effect, and an inconclusive p-value cannot prove equivalence. Effect-size language keeps those claims apart. The routing project shows how a changed pair key can reverse the analysis.
Implementation
from math import comb
def paired_sign_summary(changes):
positive = sum(change > 0 for change in changes)
negative = sum(change < 0 for change in changes)
ties = len(changes) - positive - negative
nonzero = positive + negative
if nonzero == 0:
return {"positive": 0, "negative": 0, "ties": ties, "p_two_sided": None}
smaller = min(positive, negative)
tail = sum(comb(nonzero, count) for count in range(smaller + 1))
return {"positive": positive, "negative": negative, "ties": ties,
"p_two_sided": min(1.0, 2 * tail / (2 ** nonzero))}
report = paired_sign_summary([-4, -2, 0, 1, -3])
assert report == {"positive": 1, "negative": 3, "ties": 1,
"p_two_sided": 0.625}
Performance and operating cost
Counting directions is O(n) time and O(1) extra space for n paired changes; computing the displayed binomial tail adds O(n) arithmetic. Exact arithmetic is cheap for normal study sizes. Verifying keys, timing precision and independence consumes more effort than the test itself.
Common Mistakes
- Dropping zero pairs without reporting their count or measurement precision.
- Using signed-rank location language when difference symmetry is unsupported.
- Treating repeated shifts from one agent as independent sign trials.
- Calling a nonsignificant result evidence that the two workflows are equivalent.
Read next
- Independent rank comparisons: report the pair probability, not a median claim
- Project: review routing-time differences without losing the unit
- Paired comparisons: analyze within-unit changes and preserve the match
- Paired randomization checks: enumerate sign assignments under a sharp null
- Hypothesis tests: pair the decision rule with an effect size
