Admit mature shift feedback, replay a corrected event and reject an update that forgets an older facility.
Project: release a warehouse staffing model after incremental updates
Freeze the parent and event window
The staffing forecast receives mature shift outcomes each evening. Start from one approved parent checkpoint and an immutable list of shift IDs with label revisions, feature snapshots and event times. Exclude a forecast-time feature that was actually recorded after the shift. Collapse duplicate delivery of shift 82. The admission ledger records exactly one applied revision per event.
Correct a label through replay
After the first candidate checkpoint is saved, shift 82 receives a corrected staffing outcome. Do not feed revision two as an extra gradient step on top of revision one. Return to the last clean parent, replace the event in the ordered window and replay the window into a new candidate artifact. Keep both manifests for audit. A skipped late event remains in the correction queue with an owner rather than disappearing.
Run a retention gate
The new candidate improves forecast error at a recently opened site but increases understaffing errors at a longstanding facility. Evaluate both on a pinned older-site suite and the recent mature window. Hold promotion until the older-site regression is addressed. A mean across all warehouses would conceal the harm. The release gate requires parent identity, replay integrity and quality on both frames.
Canary and hand off
For a candidate that passes, move a versioned serving pointer for one facility, record the checkpoint digest in every forecast and watch staffing overrides. Revert the pointer if the underforecast guardrail trips. Deliver applied-event watermark, excluded records, correction history, replay digest, recent and retained errors, and rollback owner. Promotion evidence closes the trace from feedback to the deployed forecast.
Implementation
def warehouse_update_window(events):
seen = {}
for event in events:
if not event["mature"] or event["feature_after_forecast"]:
return "hold:invalid-event"
previous = seen.get(event["shift_id"])
if previous is not None and previous != event["label_revision"]:
return "replay:corrected-label"
seen[event["shift_id"]] = event["label_revision"]
return "admit:" + str(len(seen))
window = [{"shift_id": "shift-47", "label_revision": "r1",
"mature": True, "feature_after_forecast": False},
{"shift_id": "shift-47", "label_revision": "r1",
"mature": True, "feature_after_forecast": False}]
assert warehouse_update_window(window) == "admit:1"
assert warehouse_update_window(window + [{**window[0],
"label_revision": "r2"}]) == "replay:corrected-label"
Performance and operating cost
The window pass is O(n) expected time and O(n) space for n events. Replaying a corrected window adds another O(u) training update pass. Retention evaluation and a canary cost extra compute, but prevent a new-site improvement from shipping with an older-site understaffing regression.
Common Mistakes
- Applying an event twice after a delivery retry.
- Training on a feature recorded after the forecast time.
- Adding a corrected label as if it were a second independent shift.
- Promoting on new-site error while an older facility gets worse.
Read next
- Incremental model updates: admit feedback once and preserve lineage
- Incremental learning releases: checkpoint replay and forgetting gates
- Promotion evidence: bind evaluation, contract and rollback to one digest
- Slice quality gates when labels are sparse or delayed
- Training checkpoints: resume state after interruption
