A count becomes comparable only after its event definition and at-risk exposure are aligned across groups.
Event counts and exposure: compare rates across unequal observation time
Choose the exposure unit
Two warehouses report equipment stoppages, but one ran many more machine-hours. Compare confirmed stoppages per operating machine-hour, not raw counts per calendar month. Exclude hours when a machine was not at risk under the declared event definition, and record changes in telemetry coverage. A rate can be scaled to a readable number of hours without changing its meaning. The target identifies which machines and operating periods are represented.
Separate counts from rates
For count C and positive exposure E, the observed rate is C divided by E. A warehouse with six stoppages over 1,200 machine-hours has the same observed rate as one with three over 600 hours. A rate ratio is meaningful only when the event rule, exposure clock and case mix are comparable. A zero event count has a zero point rate but remains uncertain, just as zero damaged packages do. The binomial interval addresses a different denominator: independent yes-or-no trials.
Inspect the Poisson working model
A simple Poisson count model assumes events arise with a stable rate over exposure and has variance equal to its mean under the model. Stoppages may cluster during a faulty shift, violating that variance pattern. Calculate counts by machine, shift and week; compare observed variation with fitted variation before using a narrow Poisson uncertainty statement. Missing machine-hours can make a warehouse appear worse even when its event count is accurate. The regression lesson extends the rate model with predictors and an exposure offset.
Publish a reviewable comparison
Show count, exposure, unit, observation period, telemetry coverage and rate together. Do not rank warehouses from a tiny count difference without uncertainty and comparable exposure. If the decision is about an intervention, account for maintenance schedules and other changes rather than treating the raw rate ratio as causal. A confounder review frames that question; the project catches a broken exposure meter.
Implementation
def machine_stoppage_rates(warehouses, per_hours=1000):
report = {}
for warehouse, record in warehouses.items():
events = record["stoppages"]
hours = record["machine_hours"]
if hours <= 0 or events < 0 or int(events) != events:
raise ValueError("invalid count or exposure")
report[warehouse] = events / hours * per_hours
return report
rates = machine_stoppage_rates({
"north": {"stoppages": 6, "machine_hours": 1200},
"south": {"stoppages": 3, "machine_hours": 600},
})
assert rates == {"north": 5.0, "south": 5.0}
Performance and operating cost
Computing rates for g groups is O(g) time and O(g) output space. Reliable at-risk exposure collection is the expensive part. A later count model can fit quickly on aggregated data, but incorrect or missing exposure can reverse the apparent ranking no matter how sophisticated the model is.
Common Mistakes
- Comparing stoppage counts without operating hours.
- Using calendar time while machines were shut down and not at risk.
- Treating a Poisson variance assumption as verified by a point rate.
- Calling an observational warehouse rate ratio the effect of a maintenance policy.
Read next
- Rare proportions: keep interval uncertainty visible at zero and one
- Project: audit rare damage and machine stoppage reports
- Poisson count models: use exposure offsets and test dispersion
- Population, estimand and sampling frame: name the quantity before calculating
- Confounders and covariate timing: adjust only for variables with a valid role
Continue the workflow: Negative binomial counts: model extra variance without inventing structural zeros.
Continue the workflow: Ratio metrics: choose ratio of totals or average of unit ratios.
