Reconstruct a paired numerator-denominator ledger, choose the business target and validate uncertainty at the depot level.
Project: audit a parcel-damage rate across uneven depots
Freeze the cohort and event clock
A transport network wants to compare a new packing rule with its prior damage rate. Define an eligible shipment, one unique parcel identifier, damage-report maturity, date attribution and depot ownership before counting. Keep zero-shipment depots in the coverage ledger even though their rates are undefined. Separate claims that arrive late from parcels whose observation window has not matured. Exposure definitions and missing outcomes help prevent a silent denominator change.
Choose the rate whose unit matches the decision
If the claim is about an average parcel, use damaged parcels over eligible parcels; if it is about a typical depot, use an explicitly weighted depot-rate estimand. Calculate both as a diagnostic when depot sizes differ. A fall in the network rate can hide worsening small depots if a large depot improves. The ratio lesson shows the arithmetic and the distinction. Preserve shipment and damage totals by depot and time period so every dashboard rate can be rebuilt.
Estimate uncertainty from independent units
Review whether treatment was assigned by depot, parcel or calendar period. Use the same unit for resampling or design variance that generated independent treatment and operational shocks. Check the largest depot’s share, positive-denominator count and the sensitivity of the result when one depot is removed for diagnosis. The delta approximation gives a transparent first pass under independent depots; it cannot rescue a four-depot network from weak independent information.
Release with explicit limits
The packet contains the data contract, maturation rule, target estimand, paired unit ledger, denominator distribution, missing-claim sensitivity and uncertainty method. The gate below holds a recommendation if the numerator and denominator windows differ or the target was never chosen. A passing packet is ready for statistical review; it does not prove the packing rule caused the observed change when the rollout was not randomized.
Implementation
def damage_rate_gate(review):
if not review["same_eligibility_window"]:
return "hold:denominator-identity"
if not review["target_unit_defined"]:
return "hold:estimand"
if not review["late_claims_accounted_for"]:
return "hold:outcome-maturity"
if not review["depot_dependence_reviewed"]:
return "hold:independent-unit"
return "review:damage-rate"
review = {"same_eligibility_window": True, "target_unit_defined": True,
"late_claims_accounted_for": False, "depot_dependence_reviewed": True}
assert damage_rate_gate(review) == "hold:outcome-maturity"
assert damage_rate_gate({**review, "late_claims_accounted_for": True}) == "review:damage-rate"
Performance and operating cost
The gate is O(1). Rebuilding the parcel ledger is O(n) expected time with unique IDs and O(n) storage for n parcels. Unit-level uncertainty then costs O(d) for d depots. The hard cost is outcome maturation, not a more elaborate division function.
Common Mistakes
- Changing the rate denominator between compared periods.
- Using a network rate to claim every depot improved.
- Treating a late claim as no damage.
- Estimating uncertainty at parcel level when the intervention and shocks operate by depot.
Read next
- Ratio metrics: choose ratio of totals or average of unit ratios
- Ratio uncertainty: retain numerator-denominator covariance
- Standard error and cluster bootstrap: resample the independent unit
- Missing outcomes: count absence before choosing an estimator
- Event counts and exposure: compare rates across unequal observation time
