A drift monitor compares serving inputs with a frozen reference distribution, while a separate matured-label cohort measures actual model error after outcomes have had time to arrive.
Feature drift and delayed-label monitoring
Watch inputs without inventing outcomes
A shipment model scores intake backlog before clearance is known. A same-day dashboard can count missing values, schema violations and backlog-bin frequencies; it cannot calculate same-day clearance error before shipments have cleared. The code compares two fixed-bin distributions with total variation distance and separately filters records whose outcomes have matured. The bins and baseline period must be frozen before inspecting the live batch.
Distinguish a warning from a failure
A high distance means the input mix changed according to this chosen summary. It does not prove predictions became worse. A new warehouse may be larger but easier to process, while a model can degrade with little marginal feature drift if the label rule changes. Investigate by site, model version and feature version. Availability at scoring time may change even when a column name does not.
Use the correct outcome clock
For clearance within 48 hours, a cohort scored yesterday is not fully labeled today. Excluding still-open shipments from an error report preferentially keeps fast clearances and biases the result. Require each scored row to pass a defined maturity window before using its outcome, then report coverage and remaining missing labels. The code intentionally excludes pending rows even when a provisional outcome field is present.
Separate alert, diagnosis and retraining
A distribution alarm opens an investigation: check upstream ingestion, missingness, site mix and timestamps. Confirm whether matured-label loss and the decision policy have changed before retraining. Retraining on the same broken feature pipeline can preserve the fault; reflexively deploying a new model also invalidates the previous comparison. The fitted pipeline and calibration checks need versioned evidence.
Define a response path
Set alert thresholds from a stable historical window and capacity for review, not from a single arbitrary number. Document owner, triage time, rollback option and the fresh validation period required for a replacement model. Monitor latency and error rate as well as statistical summaries. The release project keeps operational readiness alongside model score.
Implementation
reference_backlogs = [7, 9, 11, 14, 18, 20, 23, 27]
served_backlogs = [9, 12, 21, 24, 31, 33, 35, 39]
frozen_edges = (15, 30)
def proportions(values, edges):
if not values:
raise ValueError("empty cohort")
counts = [0] * (len(edges) + 1)
for value in values:
counts[sum(value >= edge for edge in edges)] += 1
return [count / len(values) for count in counts]
reference_mix = proportions(reference_backlogs, frozen_edges)
served_mix = proportions(served_backlogs, frozen_edges)
drift_distance = sum(abs(previous - current)
for previous, current in zip(reference_mix, served_mix)) / 2
# scored day, clearance day (or None), observed hours (or None), predicted hours
scored_shipments = [(1, 3, 7, 6), (2, 4, 8, 7),
(4, None, None, 9), (7, 8, 5, 6)]
report_day = 10
maturity_days = 3
matured = [(actual, predicted) for scored, cleared, actual, predicted
in scored_shipments
if scored + maturity_days <= report_day
and cleared is not None and cleared <= report_day
and actual is not None]
matured_mae = sum(abs(actual - predicted) for actual, predicted in matured) / len(matured)
assert 0 <= drift_distance <= 1
assert len(matured) == 3
assert matured_mae == 1Performance and operating cost
Computing B fixed-bin counts over N served rows costs O(NB) time in this direct implementation and O(B) counter memory; binary-search bins can reduce per-row cost. Matured-label scoring is O(M) for M eligible rows. Logging, retention and privacy controls can cost more than the arithmetic.
Common Mistakes
- Do not call feature drift proof of worse prediction quality.
- Do not score only already-cleared shipments when outcomes are delayed.
- Do not retrain automatically before checking upstream data and target definitions.
Read next
- Paired bootstrap intervals for model gain
- Prediction-time feature availability: reject future information before training
- Probability calibration: test whether risk scores mean what they say
- Ensemble release review project
Continue the workflow: Online updates with delayed feedback.
Continue the workflow: Covariate shift and support overlap.
Continue the workflow: Delayed bandit rewards and outcome maturity.
