A pooled graph score can hide failures on new nodes, rare relations and small connected components.
Graph model evaluation: report sparse-node and relation-specific failures
Separate task metrics
Node classification needs class-specific precision and recall, while link prediction needs a ranking metric over a declared candidate set. An accuracy figure for one task cannot stand in for the other. Report candidate construction and horizon beside every link metric. Negative sampling changes the difficulty of the ranking problem.
Slice by graph position
A high-degree lesson with many interactions may be easy to score, while a newly published lesson with no edges needs a content fallback. Break results out by node degree, age, relation type and connected component size. A single average dominated by popular nodes can justify a release that harms the exact cold-start case the feature was meant to solve.
Check retrieval before ranking
If candidate generation omits a future-positive lesson, a graph ranker never sees it. Measure candidate recall separately from ordering quality. Also measure catalog coverage and the proportion of recommendations coming from isolated or low-degree nodes. Candidate supply and graph scoring are separate stages.
Test a small slate
For a learner with two known relevant eligible lessons, a five-item list containing one gives recall at five of one half. Report the judged-set limitation and the rank of the found item. Repeat for a cold-start learner; if no relevant edge was observed, mark the metric undefined rather than inventing zero relevance.
Implementation
def judged_recall_at_k(ranked_paths, relevant_paths, limit):
if limit < 1:
raise ValueError("limit must be positive")
relevant = set(relevant_paths)
if not relevant:
return None
return len(set(ranked_paths[:limit]) & relevant) / len(relevant)Performance and operating cost
A top-K calculation costs O(K + R) time and space for R judged positives. Computing degree-sliced reports requires O(E) graph aggregation plus scored decisions; record subgroup counts and uncertainty so small slices are not overread.
Common Mistakes
- Do not mix node-classification accuracy with link-ranking quality.
- Do not report only high-degree nodes in a graph release review.
- Do not interpret an undefined no-positive recall as zero.
