A count dataset may contain zeros from a genuine no-opportunity state and zeros generated by an ordinary low-rate count process.
Excess zeros: separate no opportunity from a stochastic zero count
Define exposure before interpreting zero
A facility records device faults per machine-week. Some machines were switched off all week; others ran for many hours without a fault. Both rows have zero faults, but the first has no operating exposure and the second provides evidence about the fault rate. A positive exposure offset belongs in a count-rate model; a true zero-exposure row should not be assigned a logged offset or mistaken for an exceptionally safe machine. The offset lesson keeps counts and time at risk aligned.
Compare observed zeros with a rate model
Given a supplied Poisson mean for each exposed row, the predicted zero probability is exp of negative mean. Sum those probabilities to obtain the expected number of zero rows under that model, then compare with the observed count. The code performs only that diagnostic. More observed zeros can arise from genuine structural states, unmodeled heterogeneity, clustering, or a wrong exposure record; the gap alone does not identify a mixture component. The dispersion lesson tests a competing explanation.
Choose a mechanism before a model family
A zero-inflated count model treats some exposed rows as belonging to a latent always-zero state while the ordinary count component can also produce zeros. A hurdle model instead models zero versus positive separately, then a positive truncated count. Both can fit many zeros, but their parameters have different operational meanings. If inactive machines are observed directly, classify or exclude zero-exposure periods using the target population instead of inferring a latent class by default. The hurdle lesson separates occurrence from positive severity.
Evaluate what a decision needs
Report exposure completeness, observed and expected zero fractions, positive-count distribution, calibration by machine type, and uncertainty. A fitted model with better likelihood may still be poor for predicting the number of faults or the chance of any fault on a future machine. The project requires both zero-frequency and positive-count checks before a maintenance claim.
Implementation
from math import exp
def zero_count_diagnostic(observed_counts, predicted_means):
if not observed_counts or len(observed_counts) != len(predicted_means):
raise ValueError("aligned exposed rows required")
if any(count < 0 or int(count) != count for count in observed_counts) or any(mean <= 0 for mean in predicted_means):
raise ValueError("nonnegative counts and positive means required")
observed_zeros = sum(count == 0 for count in observed_counts)
expected_zeros = sum(exp(-mean) for mean in predicted_means)
return observed_zeros, expected_zeros
observed, expected = zero_count_diagnostic([0, 0, 1, 3], [1.0] * 4)
assert observed == 2 and 1 < expected < 2
Performance and operating cost
The diagnostic is O(n) time and O(1) extra space for n exposed rows. Fitting a mixture is more costly and can be unstable with weakly identified components. Cleaning machine uptime is often more valuable than adding a latent zero process.
Common Mistakes
- Treating a zero-exposure machine-week as a safe exposed week.
- Calling every extra zero proof of a structural always-zero class.
- Comparing models on likelihood alone while ignoring future zero and positive-count calibration.
- Logging zero exposure or silently replacing it with one operating hour.
Read next
- Negative binomial counts: model extra variance without inventing structural zeros
- Hurdle models: separate any event from the number of events once positive
- Project: select a fault-count model from exposure, zeros and burst behavior
- Poisson count models: use exposure offsets and test dispersion
- Event counts and exposure: compare rates across unequal observation time
