Run paired and repeated measurements across the operating mass range, then make a tolerance-based device decision.
Project: decide whether a handheld parcel scale can replace the dock scale
Freeze the measurement plan
A shipping team wants handheld readings to set chargeable parcel mass. Define the dock reference procedure, handheld device model, the permitted individual discrepancy and the mass bands that matter for billing. Sample parcels from ordinary operations, including heavy and awkward loads, and measure each on both devices near the same time. Log device IDs, operator, site, timestamp and order. The parcel is the unit for agreement; repeated readings of the same parcel are nested observations, not more parcels.
Audit paired error and noise
Plot handheld-minus-dock differences against paired average mass. Compute mean bias and descriptive limits in the planned operating range; inspect whether the spread or bias grows with mass. Repeat selected parcels under one controlled setup to estimate within-item variation. The agreement lesson defines individual-difference limits, while the repeatability lesson isolates short-term noise. If a device is precise but biased, calibration may be possible; if it is unstable, a constant correction cannot solve the problem.
Evaluate the decision boundary
A parcel within a few grams of a billing threshold is more sensitive to error than one far from it. Count parcels whose device disagreement could change the fee, with mass bands and device identity shown. The acceptable discrepancy must be agreed before the results are inspected. A small average bias cannot override a wide individual-difference range. The equivalence lesson explains a different inference question about average effects and margins; it does not certify every reading.
Ship an auditable packet
Keep raw pairs, repeat sequences, missing and rejected records, calibration logs, mass-band plots, tolerance failures and a decision owner. The gate below checks that sampling covers the required range and that both pairwise agreement and repeatability have been reviewed. Passing it means the evidence is ready for a human replacement decision, not that an instrument is automatically certified for every depot or future calibration state.
Implementation
def scale_replacement_gate(audit):
if not audit["required_mass_bands_covered"]:
return "hold:operating-range"
if not audit["paired_difference_reviewed"]:
return "hold:agreement"
if not audit["repeatability_reviewed"]:
return "hold:repeatability"
if audit["undocumented_calibration_changes"]:
return "hold:calibration-history"
return "review:device-replacement"
audit = {"required_mass_bands_covered": True, "paired_difference_reviewed": True,
"repeatability_reviewed": False, "undocumented_calibration_changes": False}
assert scale_replacement_gate(audit) == "hold:repeatability"
assert scale_replacement_gate({**audit, "repeatability_reviewed": True}) == "review:device-replacement"
Performance and operating cost
The gate is O(1); reconciling and plotting n paired records is O(n) time and space. Multiple devices and operators increase the number of conditions to sample. Skipping those conditions saves collection time by narrowing the claim, not by proving broad agreement.
Common Mistakes
- Approving replacement because correlation is high.
- Reporting only mean bias when individual billing decisions depend on tails.
- Counting repeats from one parcel as independent parcel coverage.
- Changing the acceptable error after seeing the differences.
Read next
- Paired device agreement: bias, limits and decision tolerance
- Repeatability: pool within-item variation without hiding drift
- Equivalence margins: require the whole interval to fit
- Paired comparisons: analyze within-unit changes and preserve the match
- Process charts: test extra variation before blaming a weekly signal
Continue the workflow: Project: qualify a pouch-filling line against fixed weight limits.
