A negative binomial mean-variance relationship allows fault counts to vary more than a Poisson model with the same conditional mean.
Negative binomial counts: model extra variance without inventing structural zeros
Look at spread at comparable exposure
Two machines can run for similar hours and still have very different fault counts because their age, maintenance history, or working conditions differ. A conditional Poisson model has variance equal to its mean. A negative binomial model in one common parameterization has variance equal to mean plus alpha times mean squared, with nonnegative alpha. The code below calculates this variance for supplied parameters; it does not estimate alpha. The Poisson regression lesson explains where the baseline assumption enters.
Distinguish heterogeneity from a data defect
A large count spread can reflect omitted predictors, route or machine clustering, changing observation windows, coding mistakes, or real latent rate heterogeneity. Check those explanations before adding a distributional parameter. Negative binomial dispersion can produce more zeros than Poisson without a separate always-zero class. A zero-inflated model therefore needs a substantive mechanism, not just a zero histogram. The zero lesson separates an inactive device from a functioning device with zero observed faults.
Preserve rate meaning
Use the same fault definition and operating exposure in both candidate models. With a log link, the exposure offset has coefficient one by design; if machine-hours were recorded differently by site, neither distribution repairs that measurement problem. Compare conditional mean calibration, variance, zero frequency, and tail frequency on held-out machine-weeks. A model that matches total faults but badly misses the rare high-fault machines may fail the maintenance decision.
Report the inference limit
Show estimated mean, alpha and uncertainty, exposure distribution, cluster structure, and whether predictive performance improves in the groups that matter. A negative binomial fit is still an association unless assignment or causal assumptions support more. The hurdle lesson offers a distinct two-part question; the project compares them only after the exposure ledger is trustworthy.
Implementation
def negative_binomial_variance(conditional_mean, dispersion_alpha):
if conditional_mean < 0 or dispersion_alpha < 0:
raise ValueError("mean and dispersion must be nonnegative")
return conditional_mean + dispersion_alpha * conditional_mean ** 2
assert negative_binomial_variance(3, 0) == 3
assert negative_binomial_variance(3, 0.4) == 6.6
Performance and operating cost
One variance calculation is O(1). Estimating model coefficients and dispersion needs numerical optimization, and validating tails needs enough independent exposed units. An extra variance parameter is cheaper than repairing an incorrect uptime ledger but cannot substitute for it.
Common Mistakes
- Reading extra zeros alone as proof a zero-inflated model is needed.
- Ignoring different machine-hours when comparing raw counts.
- Assuming a dispersion parameter fixes omitted predictors or shared site shocks.
- Reporting a model improvement without inspecting high-fault tail predictions.
Read next
- Excess zeros: separate no opportunity from a stochastic zero count
- Hurdle models: separate any event from the number of events once positive
- Project: select a fault-count model from exposure, zeros and burst behavior
- Poisson count models: use exposure offsets and test dispersion
- Standard error and cluster bootstrap: resample the independent unit
