Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Logistic regression, log loss and odds

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Binary logistic regression maps a linear feature score to a probability and fits its coefficients by reducing log loss on labeled training rows.

Name the event before fitting

A dispatch team predicts whether a shipment will miss its handoff window. The positive class is a missed handoff, defined using a fixed outcome clock; backlog measured at intake is one possible predictor. Logistic regression models log odds as an intercept plus weighted features, then turns that score into a value between zero and one. It is a classification model despite its name. Feature timing determines whether the score can be served.

Use log loss to train the probability score

For a labeled row, log loss rises sharply when the model assigns tiny probability to an event that occurs. The code implements a one-feature batch gradient step and checks that the fitted probability rises with backlog. It keeps the calculation inspectable, but lacks a convergence check, a production optimizer, missing-value handling and regularization. Use a tested solver for deployed multi-feature models, especially when classes are rare or predictors are correlated.

Interpret odds without overclaiming

A one-unit feature increase adds its coefficient to the log odds, holding the other modeled inputs fixed. Exponentiating that coefficient gives a modeled odds multiplier, not a causal effect. If backlog and staffing policy respond to each other, the coefficient can shift when the data window changes. Scale features using training rows only, store the fitted transform, and evaluate errors by site. The preprocessing guide covers that state.

Separate ranking, calibration and action

A classifier can rank missed handoffs well while its probabilities are too high for a cost calculation. Check calibration on independent rows before interpreting 0.7 as a frequency. Then choose a decision threshold from false-alarm and missed-handoff costs; the optimizer does not learn that operational threshold automatically. The threshold lesson shows the separate decision layer.

Test against real alternatives

Compare a training-rate constant, constrained tree and linear logistic model on the same later period. Report prevalence, precision, recall, log loss and calibration support. A good average log loss may hide a failure at a small warehouse or during overnight intake. Rare-event metrics make the class mix visible. Freeze the later test before selecting feature transforms or penalty strength.

Implementation

python
from math import exp, log

training_backlogs = [0.2, 0.5, 0.9, 1.2, 1.6, 2.0, 2.4, 2.8]
missed_handoff = [0, 0, 0, 0, 1, 0, 1, 1]

def probability(score):
    if score >= 0:
        return 1 / (1 + exp(-score))
    scaled = exp(score)
    return scaled / (1 + scaled)

def log_loss(backlogs, labels, intercept, slope):
    terms = []
    for backlog, label in zip(backlogs, labels):
        chance = min(max(probability(intercept + slope * backlog), 1e-12), 1 - 1e-12)
        terms.append(-label * log(chance) - (1 - label) * log(1 - chance))
    return sum(terms) / len(terms)

intercept, slope = 0.0, 0.0
initial_loss = log_loss(training_backlogs, missed_handoff, intercept, slope)
for _ in range(1200):
    residuals = [probability(intercept + slope * backlog) - label
                 for backlog, label in zip(training_backlogs, missed_handoff)]
    intercept -= 0.08 * sum(residuals) / len(residuals)
    slope -= 0.08 * sum(backlog * residual for backlog, residual
                        in zip(training_backlogs, residuals)) / len(residuals)

assert log_loss(training_backlogs, missed_handoff, intercept, slope) < initial_loss
assert probability(intercept + slope * 0.5) < probability(intercept + slope * 2.5)

Performance and operating cost

For I full-batch iterations, N rows and F features, direct gradient updates take O(INF) time and O(N + F) working memory if residuals are retained. Serving one score is O(F). Reliable solvers add convergence controls and may have different costs for sparse or high-dimensional inputs.

Common Mistakes

  • Do not treat a fitted coefficient as a causal effect.
  • Do not use the classification threshold as a substitute for calibration.
  • Do not fit scaling or select penalties on the final test period.

Read next

Continue the workflow: Naive Bayes smoothing and dependent features.

Continue the workflow: Frozen embeddings and a linear probe.

ai-data
machine-learning
Storage details