Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Offline recommendation evaluation: replay catalog and learner state in time

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

An offline score is interpretable only for items that were eligible and observable under the logged policy.

Split by decision time

Choose a cutoff, train on earlier interactions and score later recommendation decisions. Freeze learner history, item metadata and publication state at each origin. Randomly splitting clicks can leak a learner’s future completion into past predictions. Temporal feature rules apply to recommenders as much as forecasts.

Measure stages separately

Candidate recall asks whether later-positive eligible lessons entered the pool. Ranking measures such as recall at five assess order within that pool. Catalog coverage and new-item coverage show whether a policy repeatedly promotes only a small set. Keep a no-history cohort separate because its fallback differs from the learned route.

Respect missing judgments

A held-out click is a known positive, but an unclicked or unexposed lesson is not necessarily irrelevant. Offline top-k metrics can reward the previous exposure pattern. Report that limit and run a controlled online experiment when a serving change matters. Assignment design is needed to estimate the effect of the new policy.

Audit a fixture

For a learner with two known relevant lessons, a five-item slate containing one of them has recall at five of one half under the declared judged set. Verify that a lesson published after the decision is excluded from both candidate pool and metric denominator. Count how many evaluated decisions have no judged eligible positive.

Implementation

python
def recall_at_k(ranked_paths, relevant_paths, k):
    if k < 1:
        raise ValueError("k must be positive")
    relevant = set(relevant_paths)
    if not relevant:
        return None
    return len(set(ranked_paths[:k]) & relevant) / len(relevant)

Performance and operating cost

A top-k recall computation costs O(k + R) time and space for R known relevant items. Replaying N historical decisions can dominate because each needs a time-correct eligibility snapshot; store compact catalog versions to keep evaluation reproducible.

Common Mistakes

  • Do not random-split one learner’s time-ordered events.
  • Do not label every unobserved item irrelevant.
  • Do not report only ranking metrics when candidate recall or catalog coverage is poor.

Read next

Continue the workflow: NDCG at K with a declared query denominator.

ai-data
recommendation-systems
Storage details