A first-order ratio standard error depends on how each unit’s numerator moves with its denominator, not on numerator variability alone.
Ratio uncertainty: retain numerator-denominator covariance
Linearize the paired unit
For independent depots with eligible damaged counts Y and shipment counts X, the aggregate rate is sum Y divided by sum X. An approximate large-sample standard error uses residual contributions Y minus estimated rate times X, their sample standard deviation, and the mean X. This keeps the covariance induced by a high-volume depot having more possible damage. The code implements that one-step approximation for positive aggregate exposure. It is not a small-sample exact interval and it does not turn unequal depots into equally informative parcels.
Know when the approximation breaks
With only a few depots, one depot carrying most shipments, or a denominator total near zero, a symmetric interval from the approximate standard error can be unstable or even extend below zero. Inspect the contribution of each depot and consider a design-matched resampling or an interval built for the actual count process. Never bootstrap individual parcels when entire depots share operational conditions. The cluster bootstrap states the resampling unit; the rare-proportion lesson covers a different independent-Bernoulli design.
Match the sampling design
The code assumes independent, similarly sampled unit pairs. Stratified depots, unequal inclusion chances and repeated depot-days change the variance calculation. The target ratio also matters: if leadership wants a typical depot rather than a typical parcel, an interval for a mean of depot rates follows a different estimator. The target lesson fixes that choice before any uncertainty calculation. If damage reporting is incomplete, a tight standard error describes only the observed ledger, not hidden damage.
Report the denominator risk
Show unit count, total exposure, largest unit share, zero-exposure units, the estimated rate and the method for uncertainty. Plot residual contributions against shipment volume to see whether one site controls the estimate. A normal approximation can be useful for a large, well-spread collection of independent depots; it should be accompanied by a separate sensitivity analysis when the denominators are highly uneven. The project requires that sensitivity before a performance claim.
Implementation
from math import sqrt
from statistics import mean, stdev
def paired_ratio_se(damaged_counts, shipment_counts):
if len(damaged_counts) != len(shipment_counts) or len(shipment_counts) < 2:
raise ValueError("at least two aligned depots required")
if any(shipped < 0 or not 0 <= damaged <= shipped
for damaged, shipped in zip(damaged_counts, shipment_counts)):
raise ValueError("valid aligned counts required")
mean_exposure = mean(shipment_counts)
if mean_exposure <= 0:
raise ValueError("positive average exposure required")
rate = sum(damaged_counts) / sum(shipment_counts)
contributions = [damaged - rate * shipped
for damaged, shipped in zip(damaged_counts, shipment_counts)]
return rate, stdev(contributions) / (sqrt(len(contributions)) * mean_exposure)
rate, standard_error = paired_ratio_se([3, 7, 12], [60, 300, 600])
assert round(rate, 4) == 0.0229 and standard_error > 0
Performance and operating cost
The linearized calculation is O(d) time and O(d) space for d independent units; a two-pass variance can use O(1) extra space. Wider statistical machinery is justified when the design has strata or clusters. A precise formula applied to the wrong independence unit remains misleading.
Common Mistakes
- Computing numerator and denominator standard errors separately while ignoring covariance.
- Using a symmetric ratio interval with a tiny or unstable denominator.
- Resampling parcels while the independent units are depots.
- Calling a selected-depot ratio a population rate without accounting for selection.
Read next
- Ratio metrics: choose ratio of totals or average of unit ratios
- Project: audit a parcel-damage rate across uneven depots
- Standard error and cluster bootstrap: resample the independent unit
- Survey uncertainty: count sampled clusters, strata and weight concentration
- Rare proportions: keep interval uncertainty visible at zero and one
