Audit machine-hour exposure, compare count distributions, and publish only the predictive quantities maintenance can act on.
Project: select a fault-count model from exposure, zeros and burst behavior
Freeze the machine-week frame
A factory wants to forecast maintenance work from logged device faults. Define eligible machine-weeks, powered-on hours, fault deduplication, site and machine type, and the outcome maturity cutoff. Separate powered-off weeks from active weeks with no faults. If a device was down and its fault log was unavailable, mark missing observation rather than zero. The exposure lesson supplies the denominator; the missingness lesson handles absent logs.
Diagnose before fitting
Compare observed zeros with Poisson expectations for exposed rows, inspect count variance among comparable exposure bands, and examine high-fault machine-weeks. Check whether site or device age explains spread before considering latent classes. The simple Poisson, negative binomial, zero-inflated and hurdle candidates answer different data stories; candidate selection must reflect the maintenance decision. Zero diagnostics, extra variance and two-part modeling connect to these checks.
Validate on a later period
For each model, predict chance of any fault, mean count per machine-week and probability of a costly burst on a later time window. Report calibration by site and machine type with denominators. A likelihood advantage on the training set is insufficient if a simpler count model predicts the maintenance inventory better. Preserve exposure and model versions so a logging update cannot be mistaken for a risk shift.
Gate the operational packet
The packet contains uptime reconciliation, zero-exposure rules, missing-log counts, candidate assumptions, later-period calibration, uncertainty and a staffing loss measure. The gate below refuses a published recommendation if inactive machines were counted as safe exposed weeks or burst validation is missing. Passing it permits operations review; it does not establish that changing maintenance will causally change the observed fault rate.
Implementation
def device_fault_gate(audit):
if audit["zero_exposure_as_safe"]:
return "hold:exposure-definition"
if audit["missing_logs_as_zero"]:
return "hold:missing-faults"
if not audit["later_window_validated"]:
return "hold:forecast-validation"
if not audit["burst_calibration_reviewed"]:
return "hold:high-count-tail"
return "review:maintenance-model"
audit = {"zero_exposure_as_safe": True, "missing_logs_as_zero": False,
"later_window_validated": True, "burst_calibration_reviewed": True}
assert device_fault_gate(audit) == "hold:exposure-definition"
assert device_fault_gate({**audit, "zero_exposure_as_safe": False}) == "review:maintenance-model"
Performance and operating cost
The gate is O(1). Frame reconstruction and group summaries are O(n) over machine-weeks, while fitting several count models and validating rare bursts take more time. Wrong uptime exposure is a data fault that no model comparison can cheaply fix.
Common Mistakes
- Treating powered-off machines as exposed zero-fault evidence.
- Replacing unavailable fault logs with zero counts.
- Choosing the most complex count model from training likelihood alone.
- Reporting total fault accuracy without checking burst-week calibration.
Read next
- Excess zeros: separate no opportunity from a stochastic zero count
- Negative binomial counts: model extra variance without inventing structural zeros
- Hurdle models: separate any event from the number of events once positive
- Event counts and exposure: compare rates across unequal observation time
- Missing outcomes: count absence before choosing an estimator
