An offline score is interpretable only for items that were eligible and observable under the logged policy.
Offline recommendation evaluation: replay catalog and learner state in time
Split by decision time
Choose a cutoff, train on earlier interactions and score later recommendation decisions. Freeze learner history, item metadata and publication state at each origin. Randomly splitting clicks can leak a learner’s future completion into past predictions. Temporal feature rules apply to recommenders as much as forecasts.
Measure stages separately
Candidate recall asks whether later-positive eligible lessons entered the pool. Ranking measures such as recall at five assess order within that pool. Catalog coverage and new-item coverage show whether a policy repeatedly promotes only a small set. Keep a no-history cohort separate because its fallback differs from the learned route.
Respect missing judgments
A held-out click is a known positive, but an unclicked or unexposed lesson is not necessarily irrelevant. Offline top-k metrics can reward the previous exposure pattern. Report that limit and run a controlled online experiment when a serving change matters. Assignment design is needed to estimate the effect of the new policy.
Audit a fixture
For a learner with two known relevant lessons, a five-item slate containing one of them has recall at five of one half under the declared judged set. Verify that a lesson published after the decision is excluded from both candidate pool and metric denominator. Count how many evaluated decisions have no judged eligible positive.
Implementation
def recall_at_k(ranked_paths, relevant_paths, k):
if k < 1:
raise ValueError("k must be positive")
relevant = set(relevant_paths)
if not relevant:
return None
return len(set(ranked_paths[:k]) & relevant) / len(relevant)Performance and operating cost
A top-k recall computation costs O(k + R) time and space for R known relevant items. Replaying N historical decisions can dominate because each needs a time-correct eligibility snapshot; store compact catalog versions to keep evaluation reproducible.
Common Mistakes
- Do not random-split one learner’s time-ordered events.
- Do not label every unobserved item irrelevant.
- Do not report only ranking metrics when candidate recall or catalog coverage is poor.
Read next
- Implicit feedback: distinguish preference from what the system exposed
- Project: build a guarded next-lesson recommender
- Candidate retrieval: separate broad discovery from hard eligibility
- Randomized assignment: check balance and retain every assigned unit
Continue the workflow: NDCG at K with a declared query denominator.
