A corrected outcome changes an evaluation result; retaining revision history makes the change explainable and reproducible.
Label corrections: version outcomes before rebuilding quality metrics
Keep a ledger, not a mutable label cell
A receipt initially marked fraudulent may later be cleared. Store each label revision with prediction ID, event time, recorded time, reviewer or source class, reason and superseded revision. A current view can select the latest approved revision, while a historical report should state the revision cutoff it used. Prediction-outcome joins need this selection rule before calculating a confusion matrix.
Separate real changes from processing errors
A correction can reflect new evidence, a clerical fix, a taxonomy change or a bad join. These cases have different implications. A taxonomy change may require recomputing the full evaluation cohort; a bad join may require withdrawing a prior report. Do not trigger retraining automatically because one metric moved after a data repair. Retraining decisions should examine the corrected cohort and compare the same frozen population.
Recompute dependent evidence
When an approved label revision changes, identify which evaluation windows, slice dashboards and model-comparison reports depend on that prediction. Mark stale reports, recompute them under their original cohort definitions and show the delta with counts. Avoid silently replacing a published metric, since an operator may need to explain why yesterday’s alert disappeared. The report should carry model digest, cohort cutoff, label revision cutoff and metric-code revision.
Control conflicts and access
Two reviewers may submit conflicting labels. Require an adjudication state and leave the prediction out of final quality scoring until it resolves, rather than choosing the latest timestamp blindly. Restrict sensitive adjudication notes; aggregate metrics need only label, scope and revision identity. The feedback project tests conflicting revisions and reproduces both old and current reports.
Implementation
def active_labels(revisions, cutoff):
selected = {}
for row in revisions:
if row["recorded_at"] > cutoff or row["state"] != "approved":
continue
key = row["prediction_id"]
previous = selected.get(key)
if previous is None or row["revision"] > previous["revision"]:
selected[key] = row
return {key: row["label"] for key, row in selected.items()}
history = [
{"prediction_id": "decision-47", "revision": 1,
"recorded_at": 3, "state": "approved", "label": "fraud"},
{"prediction_id": "decision-47", "revision": 2,
"recorded_at": 8, "state": "approved", "label": "clear"},
{"prediction_id": "decision-82", "revision": 1,
"recorded_at": 7, "state": "disputed", "label": "fraud"},
]
assert active_labels(history, 5) == {"decision-47": "fraud"}
assert active_labels(history, 9) == {"decision-47": "clear"}
Performance and operating cost
Scanning r label revisions takes O(r) expected time and O(u) space for u unique prediction IDs. An indexed ledger can query only revised decisions for an incremental rebuild. This compact selector assumes strictly increasing revision numbers per prediction; a production ledger must enforce uniqueness and resolve conflicting writers transactionally.
Common Mistakes
- Overwriting a label and losing the prior report’s evidence.
- Choosing the newest conflicting label without adjudication.
- Changing a dashboard number without identifying the report revision.
- Triggering retraining on a correction caused by a broken join.
Read next
- Prediction-outcome joins: evaluate only mature, matched decisions
- Project: build a versioned receipt outcome feedback loop
- Retraining decisions: require a reason and a challenger comparison
- Model monitoring: separate input drift, data faults and delayed outcomes
- Label drift: delayed outcomes and sampled quality audits
Continue the workflow: Human overrides: keep decisions, reasons and labels distinct.
Continue the workflow: Label taxonomy migrations: preserve meaning across retraining.
Continue the workflow: Project: retire a spent payment-risk holdout and qualify its successor.
Continue the workflow: Incremental model updates: admit feedback once and preserve lineage.
