A hurdle model splits the probability of crossing zero from the distribution of strictly positive counts.
Hurdle models: separate any event from the number of events once positive
Ask two operational questions
A maintenance team may need the probability that a machine faults at all this week and, separately, the expected number of faults among machine-weeks with at least one. A hurdle model has a binary occurrence component and a positive-count component that is truncated at zero. Its unconditional mean is occurrence probability times the mean among positive counts. The code calculates only that identity from supplied quantities. The excess-zero lesson explains why inactive devices should first be identified from exposure records.
Do not call the parts independent causes
The two components are a statistical factorization. They do not prove there are two physical mechanisms or that preventing the first fault has no effect on later faults. A service change could shift occurrence, positive severity, or both. If the business decision is spare-parts inventory, a total mean may matter; if it is preventive inspection scheduling, the chance of any fault may be more useful. Report each component on its own scale.
Handle exposure and truncation correctly
A positive-count model must assign no probability to zero and be evaluated against only positive observed counts; fitting an ordinary count likelihood to just positive rows without accounting for truncation is not the same model. Exposure hours can affect both the occurrence probability and positive count, but their relationships need not be identical. Compare predictions by machine type and hours, including zero fraction and upper tail, on a later validation period. The dispersion lesson gives an alternative for broad count variation.
Communicate uncertainty and deployment
Report the occurrence model, positive-count model, combined mean, each interval or uncertainty band, and the fraction of rows with incomplete exposure. A zero-heavy model is unsafe for maintenance staffing if it matches average faults but misses burst weeks. The project rejects a model chosen solely because its overall likelihood is higher and requires both operational components to calibrate.
Implementation
def hurdle_expected_faults(any_fault_probability,
mean_faults_given_positive):
if not 0 <= any_fault_probability <= 1 or mean_faults_given_positive < 1:
raise ValueError("valid occurrence probability and positive count mean required")
return any_fault_probability * mean_faults_given_positive
assert hurdle_expected_faults(0.25, 2.8) == 0.7
assert hurdle_expected_faults(0, 3.2) == 0
Performance and operating cost
Combining supplied components is O(1), while fitting occurrence and zero-truncated count models costs separate optimization. Assessing both stages and their joint uncertainty is more work than fitting one mean, but it matches two distinct maintenance decisions.
Common Mistakes
- Fitting an untruncated positive-count likelihood after deleting zeros.
- Interpreting a statistical hurdle as proof of two physical fault causes.
- Ignoring operating exposure in either component.
- Selecting a model on total mean accuracy while it misses occurrence or burst counts.
Read next
- Excess zeros: separate no opportunity from a stochastic zero count
- Negative binomial counts: model extra variance without inventing structural zeros
- Project: select a fault-count model from exposure, zeros and burst behavior
- Poisson count models: use exposure offsets and test dispersion
- Diagnostic performance: separate sensitivity, specificity and predictive value
