Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: build a versioned receipt outcome feedback loop

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Join production decisions to delayed outcomes, preserve corrections and publish only reproducible quality reports.

Define the decision ledger

Record each receipt decision with unique prediction ID, model digest, decision time, eligibility class and a privacy-safe join key. Do not treat a later rescore as the same prediction. Store outcome arrivals separately, with observed time and label revision. Mature joins decide which predictions enter a report; recent pending cases remain visible in coverage counts.

Build a frozen evaluation cohort

Choose a prediction window and a 23-day maturity rule before observing model performance. Include a low-risk sample that was never sent to manual review, plus reviewed high-risk cases, and retain the selection probability if weighted estimates will be used. Add orphan outcomes, duplicate IDs, late labels, one corrected label and one unresolved dispute. Report cohort size, observed-label coverage and missingness by model and traffic slice.

Version quality evidence

Create a report keyed by model digest, cohort window, data snapshot, label revision cutoff and metric-code revision. After an approved correction, mark the old report superseded and calculate a new report under the same cohort rules. The label ledger explains why the score changed. Do not let a correction quietly rewrite an earlier operational decision without an audit event.

Gate the next model decision

Use the mature cohort to compare incumbent and challenger with the same sampling and label rules. If coverage is low or reviewer selection is biased, mark the quality result inconclusive. A drift alert may trigger investigation, but only sufficient, reproducible outcome evidence should support a retraining or promotion decision. Challenger evaluation and promotion gates consume this report.

Implementation

python
def feedback_report(predictions, resolved_labels, minimum_coverage=0.72):
    if not predictions:
        return {"state": "hold", "reason": "empty-cohort"}
    ids = [row["prediction_id"] for row in predictions]
    if len(ids) != len(set(ids)):
        return {"state": "hold", "reason": "duplicate-prediction"}
    matched = [row for row in predictions
               if row["prediction_id"] in resolved_labels]
    coverage = len(matched) / len(predictions)
    if coverage < minimum_coverage:
        return {"state": "inconclusive", "coverage": coverage}
    correct = sum(row["predicted_label"] ==
                  resolved_labels[row["prediction_id"]] for row in matched)
    return {"state": "reportable", "coverage": coverage,
            "accuracy_on_matched": correct / len(matched)}

cohort = [{"prediction_id": "decision-47", "predicted_label": "fraud"},
          {"prediction_id": "decision-82", "predicted_label": "clear"}]
assert feedback_report(cohort, {"decision-47": "fraud"})["state"] == "inconclusive"
assert feedback_report(cohort, {"decision-47": "fraud",
                                "decision-82": "clear"})["state"] == "reportable"

Performance and operating cost

The compact report scans p predictions in O(p) expected time and uses O(p) space for duplicate detection and matches. Joining revisions, calculating slices and adjusting for known sampling probabilities add work. Accuracy on matched cases is not population accuracy when label availability depends on the model decision; retain that limitation in every report.

Common Mistakes

  • Publishing accuracy from an immature recent cohort.
  • Collapsing rescored decisions under one receipt ID.
  • Hiding low label coverage behind a precise percentage.
  • Editing an old report in place after a label correction.

Read next

ai-data
mlops
Storage details