Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Online updates with delayed feedback

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An online classifier updates its parameters from newly matured labels in event order, after it has recorded each prediction, rather than rebuilding from the entire history on every arrival.

Score before learning from the event

A handoff model predicts at intake. The actual missed-handoff label arrives later, sometimes after a route investigation. Record the prediction, model version and feature snapshot first. Only a matured label can trigger an update. The code simulates chronological days, scores arriving shipments, and applies small logistic-gradient steps for earlier cases whose labels became available. The event’s own hidden label is never read during its prediction.

Keep a fixed feature contract

Online updating does not remove preprocessing risk. Store feature order, scaling state, missing-value policy and label definition with each model version. Recomputing a global mean from future rows while replaying historical examples changes old inputs. The pipeline guide defines the fitted transform contract; feature timing defines the input clock.

Evaluate predict-then-update honestly

A streaming report can compare each stored prediction with its label after maturity. It must not re-score the event using a model already updated from that event. Report the count of mature outcomes and the backlog of pending labels. Early reports can overrepresent fast resolutions. Delayed-label monitoring explains that censoring problem.

Watch drift and forgetting

A fixed learning rate can chase recent noise, while a decaying rate can stop adapting after a real process change. Keep a replay sample from earlier periods or compare with a periodically retrained model if older regimes still matter. Do not update solely because an input histogram changed; confirm feature integrity and mature-label loss. An online fit can be cheaper per update but harder to reproduce and roll back than a batch model.

Use a release gate for parameter changes

Log update order, label arrival, learning rate, seed if sampling and an immutable checkpoint. Evaluate proposed update rules on a historical stream without accessing labels before their arrival. Keep a sealed later stream for a final comparison. The applied review brings the update and monitoring decisions together.

Implementation

python
from math import exp

# Arrival day, shipment ID, backlog, label day, eventual missed-handoff label.
arrivals = [(1, "S501", 0.4, 3, 0), (2, "S502", 1.8, 5, 1),
            (4, "S503", 1.2, 7, 0), (6, "S504", 2.3, 8, 1)]
intercept, slope = 0.0, 0.0
pending = []
predictions = {}
learning_rate = 0.15

def probability(backlog, fitted_intercept, fitted_slope):
    score = fitted_intercept + fitted_slope * backlog
    return 1 / (1 + exp(-score))

for day in range(1, 10):
    for arrival_day, shipment_id, backlog, label_day, missed in arrivals:
        if arrival_day == day:
            predictions[shipment_id] = probability(backlog, intercept, slope)
            pending.append((shipment_id, backlog, label_day, missed))
    matured = [record for record in pending if record[2] == day]
    pending = [record for record in pending if record[2] > day]
    for shipment_id, backlog, label_day, missed in matured:
        residual = probability(backlog, intercept, slope) - missed
        intercept -= learning_rate * residual
        slope -= learning_rate * backlog * residual

assert len(predictions) == 4
assert not pending
assert predictions["S501"] == 0.5
assert all(0 < score < 1 for score in predictions.values())

Performance and operating cost

Each one-feature event score and mature-label update costs O(1) arithmetic and parameter storage; with F features the update is O(F). A pending-label queue uses O(P) space for P unresolved cases. Logging, replay buffers, checkpoints and late correction handling add operational cost.

Common Mistakes

  • Do not use a shipment’s outcome before its label arrival day.
  • Do not report re-scored historical predictions as pre-update accuracy.
  • Do not change scaling or target definitions silently between checkpoints.

Read next

Continue the workflow: Client participation, dropouts and stale updates.

ai-data
machine-learning
Storage details