Join production decisions to delayed outcomes, preserve corrections and publish only reproducible quality reports.
Project: build a versioned receipt outcome feedback loop
Define the decision ledger
Record each receipt decision with unique prediction ID, model digest, decision time, eligibility class and a privacy-safe join key. Do not treat a later rescore as the same prediction. Store outcome arrivals separately, with observed time and label revision. Mature joins decide which predictions enter a report; recent pending cases remain visible in coverage counts.
Build a frozen evaluation cohort
Choose a prediction window and a 23-day maturity rule before observing model performance. Include a low-risk sample that was never sent to manual review, plus reviewed high-risk cases, and retain the selection probability if weighted estimates will be used. Add orphan outcomes, duplicate IDs, late labels, one corrected label and one unresolved dispute. Report cohort size, observed-label coverage and missingness by model and traffic slice.
Version quality evidence
Create a report keyed by model digest, cohort window, data snapshot, label revision cutoff and metric-code revision. After an approved correction, mark the old report superseded and calculate a new report under the same cohort rules. The label ledger explains why the score changed. Do not let a correction quietly rewrite an earlier operational decision without an audit event.
Gate the next model decision
Use the mature cohort to compare incumbent and challenger with the same sampling and label rules. If coverage is low or reviewer selection is biased, mark the quality result inconclusive. A drift alert may trigger investigation, but only sufficient, reproducible outcome evidence should support a retraining or promotion decision. Challenger evaluation and promotion gates consume this report.
Implementation
def feedback_report(predictions, resolved_labels, minimum_coverage=0.72):
if not predictions:
return {"state": "hold", "reason": "empty-cohort"}
ids = [row["prediction_id"] for row in predictions]
if len(ids) != len(set(ids)):
return {"state": "hold", "reason": "duplicate-prediction"}
matched = [row for row in predictions
if row["prediction_id"] in resolved_labels]
coverage = len(matched) / len(predictions)
if coverage < minimum_coverage:
return {"state": "inconclusive", "coverage": coverage}
correct = sum(row["predicted_label"] ==
resolved_labels[row["prediction_id"]] for row in matched)
return {"state": "reportable", "coverage": coverage,
"accuracy_on_matched": correct / len(matched)}
cohort = [{"prediction_id": "decision-47", "predicted_label": "fraud"},
{"prediction_id": "decision-82", "predicted_label": "clear"}]
assert feedback_report(cohort, {"decision-47": "fraud"})["state"] == "inconclusive"
assert feedback_report(cohort, {"decision-47": "fraud",
"decision-82": "clear"})["state"] == "reportable"
Performance and operating cost
The compact report scans p predictions in O(p) expected time and uses O(p) space for duplicate detection and matches. Joining revisions, calculating slices and adjusting for known sampling probabilities add work. Accuracy on matched cases is not population accuracy when label availability depends on the model decision; retain that limitation in every report.
Common Mistakes
- Publishing accuracy from an immature recent cohort.
- Collapsing rescored decisions under one receipt ID.
- Hiding low label coverage behind a precise percentage.
- Editing an old report in place after a label correction.
Read next
- Prediction-outcome joins: evaluate only mature, matched decisions
- Label corrections: version outcomes before rebuilding quality metrics
- Model monitoring: separate input drift, data faults and delayed outcomes
- Retraining decisions: require a reason and a challenger comparison
- Project: release receipt triage with lineage, canary checks and rollback
