Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Paired device agreement: bias, limits and decision tolerance

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Agreement asks whether two instruments give close enough readings on the same units for the intended decision; correlation asks a different question.

Pair before calculating

A depot weighs each sealed parcel once with a dock scale and once with a handheld scale. The comparison unit is the parcel, not the individual reading. Define the difference as handheld minus dock and keep the two readings close enough in time that a change in the parcel cannot explain the gap. A strong correlation can coexist with a constant two-kilogram bias because parcels differ widely in true mass. Paired estimands explain why unpaired averages discard the central comparison.

Separate bias from spread

The mean paired difference estimates average directional bias in the sampled operating range. A useful descriptive interval for individual future differences, when independent differences are reasonably stable and approximately normal, is mean difference plus or minus 1.96 times their sample standard deviation. This is a limit of agreement, not a confidence interval for the mean and not a pass threshold chosen after seeing the data. The code calculates descriptive limits; it does not quantify uncertainty in those limits. Interval coverage addresses the separate question of repeated-sample uncertainty.

Look for proportional error

Plot each pair’s difference against its paired average. If heavier parcels have larger differences, one common bias and one common spread may mislead; report agreement by mass band or choose a justified scale before fitting a calibration. Also inspect duplicate parcel IDs, unit conversions, missing pairs and repeated readings from the same parcel. Repeated pairs are dependent unless the analysis accounts for the parcel. The repeatability lesson tests whether either instrument is noisy even when the parcel is unchanged.

Set tolerance from the operation

If a weight discrepancy can trigger a shipping charge, specify the acceptable individual difference in kilograms before evaluating devices. A narrow average-bias interval cannot certify that nearly every parcel meets that tolerance. Report pair count, range of masses, mean bias, limits, their uncertainty, outlying pairs and the fraction beyond tolerance. Agreement only covers the conditions sampled: an indoor calibration run does not establish performance on vibrating loading bays. The project turns these findings into a replacement decision.

Implementation

python
from statistics import mean, stdev

def agreement_limits(handheld_kg, dock_kg):
    if len(handheld_kg) != len(dock_kg) or len(dock_kg) < 2:
        raise ValueError("at least two aligned parcel pairs required")
    differences = [handheld - dock for handheld, dock in zip(handheld_kg, dock_kg)]
    bias = mean(differences)
    spread = stdev(differences)
    return bias, (bias - 1.96 * spread, bias + 1.96 * spread)

bias, limits = agreement_limits([12.4, 21.8, 33.1], [12.0, 22.0, 32.5])
assert round(bias, 3) == 0.267 and limits[0] < 0 < limits[1]

Performance and operating cost

Pair matching and descriptive limits take O(n) time and O(n) space for n parcels. Streaming mean and variance can reduce storage to O(1), but storing pairs supports the necessary difference-versus-average plot and record audit. More pairs do not correct a changed weighing procedure.

Common Mistakes

  • Substituting correlation or regression fit for agreement.
  • Calling the limits a confidence interval for average bias.
  • Choosing a tolerance after seeing the device differences.
  • Treating repeated readings of one parcel as independent parcels.

Read next

ai-data
applied-statistics
Storage details