A ratio of aggregate numerators to aggregate denominators answers a different question from the average of unit-level ratios.
Ratio metrics: choose ratio of totals or average of unit ratios
Define the unit and target
A courier team asks for damaged parcels per thousand shipped parcels across depots. If the target is the risk for a randomly selected parcel, sum damaged parcels across eligible depots and divide by all shipped parcels. If the target is the experience of a randomly selected depot, average depot-specific rates under a stated depot weighting scheme. These targets need not agree. A small depot with four damaged parcels in forty shipments counts as much as a huge depot under an unweighted depot average, but contributes only forty shipments to the parcel-weighted rate.
Keep numerator and denominator aligned
The code calculates both summaries for an original three-depot ledger and refuses impossible records. It excludes zero-shipment depots from the average of defined rates, while still requiring zero damage for such a depot. For a parcel-level risk, numerator and denominator must cover the same dates, eligible services and deduplication rules. A late damage claim counted against last month’s shipments can distort the rate even if every arithmetic operation is correct. Count-rate exposure explains another denominator tied to time at risk.
Do not collapse the design
If depots were sampled with unequal probabilities, summing sampled shipments is not automatically a population parcel rate. A design-based estimator needs the sampling and response weights on both totals. Repeated days within depot can share shocks, so thousands of parcel rows may not yield thousands of independent uncertainty units. Survey design and cluster resampling address those boundaries. The next lesson derives a paired-denominator approximation for independent units, then warns when it fails.
Publish both when decisions differ
A policy that improves the overall parcel-weighted rate could still leave small depots worse off. Report total damaged parcels, total eligible shipments, each depot’s rate, zero-denominator handling and the intended unit of fairness. A single headline rate cannot show whether the change is broad or driven by one high-volume depot. Ratio uncertainty covers the covariance between counts and exposure; the project ties denominator identity to a release gate.
Implementation
def depot_damage_rates(depot_rows):
if not depot_rows:
raise ValueError("at least one depot required")
total_damage = total_shipments = 0
defined_rates = []
for depot_id, damaged, shipped in depot_rows:
if not depot_id or shipped < 0 or not 0 <= damaged <= shipped:
raise ValueError("valid aligned depot counts required")
total_damage += damaged
total_shipments += shipped
if shipped:
defined_rates.append(damaged / shipped)
if not total_shipments:
raise ValueError("no eligible shipments")
return total_damage / total_shipments, sum(defined_rates) / len(defined_rates)
parcel_rate, depot_mean_rate = depot_damage_rates([
("Harbor", 2, 40), ("Ridge", 8, 400), ("Central", 14, 700)])
assert round(parcel_rate, 4) == 0.0211
assert round(depot_mean_rate, 4) == 0.03
Performance and operating cost
Scanning d depot rows costs O(d) time and O(d) space for the stored defined rates; a running sum gives O(1) extra space. Rebuilding deduplicated parcel counts can cost far more than the division. The calculation cannot repair an incompatible date window or an unrecorded zero-exposure depot.
Common Mistakes
- Calling the unweighted average depot rate a parcel-level risk.
- Including damage claims and shipments from different eligible windows.
- Replacing zero shipments with one to avoid division by zero.
- Treating all parcel rows as independent when depot shocks dominate.
Read next
- Ratio uncertainty: retain numerator-denominator covariance
- Project: audit a parcel-damage rate across uneven depots
- Event counts and exposure: compare rates across unequal observation time
- Survey uncertainty: count sampled clusters, strata and weight concentration
- Standard error and cluster bootstrap: resample the independent unit
