A categorical naive Bayes classifier combines a training-derived class prior with per-feature likelihoods while assuming features are independent after conditioning on class.
Naive Bayes smoothing and dependent features
State the conditional assumption precisely
A missed-handoff triage model might observe an intake scan gap and a staffing alert. Naive Bayes multiplies each feature likelihood conditional on whether the handoff was missed. It does not claim the features are unconditionally independent. If both alerts are generated from the same system outage, their evidence is counted twice under the model. The ranking may remain useful, but the numeric probability can be overconfident. Calibration must be checked separately.
Smooth unseen combinations
A rare class may have no training case with one alert present. Without smoothing, its likelihood becomes zero and can overwhelm all other evidence. The code applies add-one smoothing to Bernoulli feature counts, then adds log likelihoods to avoid multiplying tiny numbers. This is a small binary-feature teaching model; event counts in documents or continuous measurements require different likelihood families and assumptions.
Keep the class prior from the right period
The prior reflects the positive share in the training cohort. It may differ after a carrier contract or label policy changes. Do not insert the future test prevalence into the training prior to make a result look better. Report the observed prevalence by development and final cohorts. The rare-event guide shows why accuracy and precision shift with class mix.
Inspect correlated signals and missingness
A scan-gap alert and delayed-scan count can be almost the same event. Including both can produce extreme log odds without new information. Compare performance with one removed, inspect conditional co-occurrence and distinguish an absent alert from a missing sensor reading. Treating missing as false changes the question. Feature availability needs a serving-time contract.
Compare against a calibrated alternative
Use the same future rows to compare this model with a training-rate baseline and logistic regression. Evaluate precision, recall and log loss before deciding whether its simplicity is worth the independence approximation. A probability threshold comes from intervention cost, not from the Bayes formula itself. Decision policy remains a separate layer.
Implementation
from math import exp, log
# scan gap, staffing alert, missed handoff
training_events = [
(0, 0, 0), (0, 1, 0), (0, 0, 0), (1, 0, 0),
(1, 1, 1), (1, 0, 1), (0, 1, 1), (1, 1, 1),
]
def log_class_score(signals, missed):
matching = [event for event in training_events if event[2] == missed]
prior = len(matching) / len(training_events)
score = log(prior)
for column, present in enumerate(signals):
present_count = sum(event[column] for event in matching)
smoothed_present = (present_count + 1) / (len(matching) + 2)
score += log(smoothed_present if present else 1 - smoothed_present)
return score
def missed_probability(signals):
log_safe = log_class_score(signals, 0)
log_missed = log_class_score(signals, 1)
odds = exp(log_missed - log_safe)
return odds / (1 + odds)
assert missed_probability((1, 1)) > missed_probability((0, 0))
assert 0 < missed_probability((1, 0)) < 1Performance and operating cost
Counting class-feature pairs takes O(NF) training time and O(CF) model storage for N rows, F binary features and C classes. A score costs O(CF) time. Log arithmetic prevents many underflow failures, but correlated features and shifted priors still require evaluation.
Common Mistakes
- Do not interpret the conditional independence assumption as a fact about the process.
- Do not allow an unseen feature value to create a zero likelihood without a policy.
- Do not assume the returned probability is calibrated.
