Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Prediction-outcome joins: evaluate only mature, matched decisions

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A production quality metric needs the exact prediction, its later outcome and a declared maturity rule.

Join decisions by immutable identity

A receipt can be rescored after a correction. Joining an outcome to a receipt ID alone may attach it to the wrong decision. Store a unique prediction ID, model digest, decision timestamp, feature-contract version and decision purpose. The outcome record should identify which decision it evaluates and when the result became known. Inference logging defines a privacy-conscious join key; a raw receipt image need not be copied into the metric table.

Declare an observation window

A disputed transaction may be resolved 23 days after prediction. Recent decisions lack the same opportunity to acquire a positive outcome. Mark them pending rather than negative. A maturity delay is a policy tied to the outcome process, not an arbitrary wait that improves the chart. Track coverage: how many eligible predictions have observed outcomes and which populations remain missing. Delayed-label monitoring separates current serving health from mature quality.

Avoid outcome selection bias

Some decisions are reviewed manually because the model marked them risky. Those decisions receive labels sooner and more often than low-risk decisions. A quality estimate from reviewed cases alone describes the reviewed population. Preserve the selection rule, sample a known fraction of low-risk cases if lawful and practical, and report population-weighted or stratified results only when the sampling design supports them. An outcome feed can also contain corrections; keep versions rather than overwriting history.

Audit join failures

Count duplicate prediction IDs, orphan outcomes, multiple outcomes, late arrivals and missing coverage by model version and slice. Test a record whose outcome arrives after the cutoff and a receipt with two distinct predictions. The feedback project builds an evaluation set from this ledger and refuses to claim accuracy when the label population is too sparse.

Implementation

python
from datetime import timedelta

def mature_pairs(predictions, outcomes, cutoff, maturity_days=23):
    by_prediction = {}
    for outcome in outcomes:
        key = outcome["prediction_id"]
        if key in by_prediction:
            raise ValueError("duplicate outcome for prediction")
        by_prediction[key] = outcome
    pairs = []
    for prediction in predictions:
        if prediction["at"] > cutoff - timedelta(days=maturity_days):
            continue
        outcome = by_prediction.get(prediction["prediction_id"])
        if outcome is not None and outcome["observed_at"] <= cutoff:
            pairs.append((prediction, outcome))
    return pairs

from datetime import datetime, timezone
cutoff = datetime(2026, 10, 6, tzinfo=timezone.utc)
old = {"prediction_id": "decision-47", "at": cutoff - timedelta(days=27)}
new = {"prediction_id": "decision-82", "at": cutoff - timedelta(days=8)}
labels = [{"prediction_id": "decision-47", "observed_at": cutoff}]
assert len(mature_pairs([old, new], labels, cutoff)) == 1

Performance and operating cost

Joining p predictions to o outcomes uses O(p + o) expected time and O(o) lookup memory. Materialized joins can reduce repeated scans but need correction handling and retention limits. The example rejects duplicate outcomes; a full ledger should version revisions and choose the active label under an explicit rule.

Common Mistakes

  • Using receipt ID when one receipt has several predictions.
  • Counting pending outcomes as negative labels.
  • Measuring only cases selected for manual review as if they were all traffic.
  • Overwriting corrected outcomes without preserving the earlier version.

Read next

Continue the workflow: Batch replay: supersede outputs without duplicating downstream actions.

Continue the workflow: Model experiment guardrails: stop harm without misreading the sample.

Continue the workflow: Slice quality gates when labels are sparse or delayed.

Continue the workflow: Selective labels: measure what the model never lets reviewers see.

Continue the workflow: Calibration monitoring: pair score bands with matured outcomes.

Continue the workflow: Prediction-set monitoring: coverage lag, set width and review load.

Continue the workflow: Project: evaluate a pickup-notification bandit before rollout.

ai-data
mlops
Storage details