A linear margin classifier seeks a separating score with a buffer around the boundary, penalizing examples that fall inside that buffer or on the wrong side.
Linear margin classifiers and hinge loss
Distinguish a margin from a probability
A missed-handoff model may output a signed score: positive for predicted miss, negative otherwise. The hinge penalty for a labeled row is the amount by which its signed margin falls below one, floored at zero. This encourages a gap around the boundary. The raw score is not a calibrated chance of failure. Logistic regression optimizes a different loss and yields probabilities that still require calibration checks.
Balance the buffer against violations
A soft-margin linear classifier combines hinge loss with a penalty on weight size. Stronger weight regularization allows more margin violations but may improve stability under noisy labels. An overly weak penalty can follow mislabeled outliers. The example uses a one-dimensional batch subgradient with fixed training-scale inputs; a production solver needs convergence diagnostics and an explicit feature-scaling pipeline. Fit scaling on each training fold.
Inspect cases closest to the boundary
Rows beyond the margin do not contribute to the hinge term; rows inside or across it do. A support-vector implementation exploits this geometry, though the simple code updates against every row. Investigate borderline shipments for inconsistent target timestamps, mixed sites and missing features. A boundary can move sharply when a rare site has only a handful of labels. Prevalence and recall still matter after the margin is fitted.
Evaluate the operational policy separately
Choose a score cutoff against a development cohort according to review capacity and missed-handoff cost. If downstream dispatch needs a numeric probability, calibrate the score using disjoint data; do not label the score itself as 80-percent confidence. For a new site or future month, report ranking and decision metrics separately. The decision-policy lesson provides the cost calculation.
Know when to switch model family
A linear boundary may be too simple when backlog interacts strongly with shift staffing or route mix. Add only features available at scoring time, then compare with a constrained tree or forest. Kernels can create nonlinear boundaries but add fitting, memory and serving costs. Use development-only selection and leave the final future test sealed.
Implementation
training_rows = [(-2.0, -1), (-1.3, -1), (-0.8, -1),
(0.6, 1), (1.1, 1), (1.8, 1)]
weight, intercept = 0.0, 0.0
weight_penalty = 0.04
learning_rate = 0.12
def objective(rows, fitted_weight, fitted_intercept):
hinge = sum(max(0.0, 1 - label * (fitted_weight * backlog + fitted_intercept))
for backlog, label in rows) / len(rows)
return hinge + weight_penalty * fitted_weight ** 2 / 2
initial_objective = objective(training_rows, weight, intercept)
for _ in range(500):
violations = [(backlog, label) for backlog, label in training_rows
if label * (weight * backlog + intercept) < 1]
gradient_weight = weight_penalty * weight - sum(
label * backlog for backlog, label in violations) / len(training_rows)
gradient_intercept = -sum(label for _, label in violations) / len(training_rows)
weight -= learning_rate * gradient_weight
intercept -= learning_rate * gradient_intercept
assert objective(training_rows, weight, intercept) < initial_objective
assert all(label * (weight * backlog + intercept) > 0
for backlog, label in training_rows)Performance and operating cost
The dense full-batch teaching loop takes O(INF) time over I iterations, N rows and F features, with O(F) model storage. A kernel classifier can retain many training cases and become much more expensive to train and serve. Measure actual latency before choosing it over a linear or tree model.
Common Mistakes
- Do not interpret a margin score as a probability.
- Do not scale validation or serving features with new fitted statistics.
- Do not select the margin penalty using the final test labels.
