Class weighting changes the training loss assigned to each labeled class, while a decision threshold changes which already-produced scores trigger an action.
Class weighting versus decision thresholds
Separate the two controls
A missed handoff occurs less often than an on-time handoff. Weighting positive examples more heavily asks the optimizer to spend more capacity fitting them; moving the alert threshold changes the dispatch rule after training. These interventions can produce similar recall at one operating point but different rankings and probability behavior. Choose the operation according to the failure mode rather than treating the controls as interchangeable. The threshold guide defines the action layer.
Estimate weights from training labels only
The code sets a declared positive weight and fits two one-feature logistic models, one weighted and one unweighted. Weighted loss can shift a probability upward at the same backlog. A balanced inverse-frequency weight is one possible starting value, but its formula reflects class counts, not the real cost of a missed handoff. Do not use final-test prevalence to calculate training weights. Prevalence belongs in every evaluation report.
Recheck calibration after weighting
A weighted classifier’s numeric outputs often should not be read as event frequencies without checking them on representative, disjoint labeled data. If dispatch computes expected cost from those values, calibration matters. A lower threshold applied to a calibrated model may be easier to audit than changing the training objective; the comparison must use the same future cohort. Calibration checks expose a shifted probability scale.
Inspect false alarms and missed events
Report precision, recall, number of daily alerts, and errors by site at each candidate operating point. An alert policy that catches most misses but creates more reviews than the desk can process is not usable. Conversely, a weighted fit can help if the learner ignored a rare class entirely. Select settings inside development data, then verify once on an untouched later period. Model selection protects that claim.
Keep labels and the intervention stable
Weights cannot correct a mislabeled outcome, a new carrier mix or a feature that arrives after the alert decision. If a staffing change reduces misses, the target rate and action value can shift. Monitor outcome maturity and post-deployment error before retraining. The monitoring guide separates immediate feature changes from delayed error evidence.
Implementation
from math import exp
shipments = [(0.1, 0), (0.3, 0), (0.5, 0), (0.8, 0),
(1.1, 0), (1.4, 0), (1.7, 0), (2.0, 1)]
def chance(intercept, slope, backlog):
score = intercept + slope * backlog
return 1 / (1 + exp(-score))
def fit_weighted_logistic(positive_weight):
intercept, slope = 0.0, 0.0
for _ in range(1000):
intercept_gradient = 0.0
slope_gradient = 0.0
for backlog, missed in shipments:
weight = positive_weight if missed else 1.0
residual = weight * (chance(intercept, slope, backlog) - missed)
intercept_gradient += residual
slope_gradient += backlog * residual
intercept -= 0.06 * intercept_gradient / len(shipments)
slope -= 0.06 * slope_gradient / len(shipments)
return intercept, slope
plain = fit_weighted_logistic(1.0)
upweighted = fit_weighted_logistic(5.0)
assert chance(*upweighted, 1.5) > chance(*plain, 1.5)
assert 0 < chance(*plain, 1.5) < 1Performance and operating cost
A dense one-feature full-batch fit takes O(IN) time over I iterations and N rows; with F features it is O(INF). Applying a new threshold to stored scores is O(N) and needs no retraining. Calibrating and auditing either policy adds data and review cost.
Common Mistakes
- Do not derive training weights from final-test labels.
- Do not assume weighted output remains calibrated.
- Do not optimize recall without checking review capacity and precision.
Read next
- Rare-event precision, recall and changing prevalence
- Decision thresholds: choose an action from probabilities and error costs
- Probability calibration: test whether risk scores mean what they say
- Handoff triage model review project
Continue the workflow: Label-prior shift and odds adjustment.
