Repeatability measures variation when the same item is measured again under specified near-identical conditions.
Repeatability: pool within-item variation without hiding drift
Define the conditions
A handheld scale reads the same sealed parcel three times without changing the load, operator or setting. Differences among those readings represent short-term measurement variation under that protocol. Change operator, location, temperature or day and the question becomes broader reproducibility or stability, not short-term repeatability. Preserve those fields in the measurement table so the source of spread is visible. Agreement between devices includes their biases as well as their noise.
Pool within each item
When several parcels are repeated, compute each parcel’s mean and sum its squared deviations from that mean. Divide the combined sum by the total within-parcel degrees of freedom, then take a square root. This estimates a pooled within-item standard deviation when a common short-term variance is plausible. The implementation accepts unequal repeat counts and rejects single-reading groups; it does not estimate a between-item component. Pooling all raw readings around one grand mean would mix actual parcel mass differences into measurement noise.
Check whether one spread fits
Inspect within-parcel spread against true or paired-average mass, operator and measurement order. If the scale becomes noisier above thirty kilograms, a single pooled standard deviation can understate high-mass uncertainty. Order also matters: a slow zeroing drift may yield systematic first-to-third differences that a pooled variance hides. Randomize or alternate order where practical and retain the sequence. Process-shift analysis addresses later drift, but it cannot repair missing order metadata.
Translate spread into a decision
For two independent repeated readings from one stable instrument, their difference has standard deviation about square root of two times the within-item standard deviation under the simple constant-variance model. That statement depends on independence and the same conditions; shared calibration error does not vanish. Compare the likely repeat gap with a predeclared tolerance, and report item count, repeats per item, pooled estimate, per-item spread and order diagnostics. The project refuses a device swap if precision fails despite small average bias.
Implementation
from math import sqrt
from statistics import mean
def pooled_repeatability(parcel_repeats):
if not parcel_repeats or any(len(readings) < 2 for readings in parcel_repeats.values()):
raise ValueError("each parcel needs at least two readings")
squared_error = 0.0
degrees_of_freedom = 0
for readings in parcel_repeats.values():
parcel_mean = mean(readings)
squared_error += sum((reading - parcel_mean) ** 2 for reading in readings)
degrees_of_freedom += len(readings) - 1
return sqrt(squared_error / degrees_of_freedom)
repeat_sd = pooled_repeatability({"P47": [14.1, 14.3], "P82": [29.6, 29.4]})
assert round(repeat_sd, 3) == 0.141
Performance and operating cost
The pooled calculation is O(n) time for n readings and O(1) extra space beyond the grouped input. Storing order, operator and instrument identifiers costs O(n) but is necessary to diagnose changing precision. An extra decimal place in the output cannot compensate for a flawed repeat protocol.
Common Mistakes
- Pooling around one grand mean and calling parcel-to-parcel variation instrument noise.
- Claiming short-term repeats cover different operators or weather conditions.
- Ignoring a trend by measurement order because pooled spread is small.
- Using a pooled value when precision changes sharply by mass.
Read next
- Paired device agreement: bias, limits and decision tolerance
- Project: decide whether a handheld parcel scale can replace the dock scale
- Process shifts: separate a scheduled intervention from a searched break
- Paired comparisons: analyze within-unit changes and preserve the match
- Process charts: test extra variation before blaming a weekly signal
Continue the workflow: Process capability: check stability before interpreting Cpk.
