Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Neighborhood features: aggregate only the graph that existed at prediction time

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Message passing updates a node representation from its neighbors; the neighborhood and attributes must obey the same time boundary as the target.

Start with a baseline

Before training a graph neural network, count eligible prerequisite neighbors or shared concepts and compare that simple score with a content baseline. A learned message-passing layer combines node features with neighbor information, but a two-hop path can already reveal useful curriculum structure. Content features offer a fallback when learner edges are sparse.

Control the receptive field

One layer sees immediate neighbors, two layers can see neighbors of neighbors. High-degree concept nodes make multi-hop expansion grow rapidly. Choose layer count and sampling fanout by the question, not by an arbitrary large network. Track which relation types carry messages and whether a learner can receive information from a future completion edge.

Avoid target leakage

For link prediction, remove the target learner-lesson edge from the message-passing graph before scoring that pair. Otherwise the model can observe the answer directly. Temporal splits also require later edges to be absent from earlier graph snapshots. Link evaluation must freeze both graph structure and candidate catalog.

Test propagation

Use a lesson connected to two concept nodes with numeric importance values 2 and 6. Its mean neighbor value is 4. Add a duplicate of the value-6 edge and verify that unique-neighbor aggregation still returns 4. Then move one edge after the prediction time and check the earlier feature changes under the snapshot rule.

Implementation

python
def mean_neighbor_feature(neighbor_ids, feature_by_node):
    unique_ids = set(neighbor_ids)
    values = [feature_by_node[node_id] for node_id in unique_ids if node_id in feature_by_node]
    return None if not values else sum(values) / len(values)

Performance and operating cost

A one-hop aggregation is O(D) for D neighbors; L layers with fanout F can visit up to O(F to the power L) sampled positions per seed before deduplication. Measure actual sampled nodes and memory rather than assuming a deeper network is free.

Common Mistakes

  • Do not leave a target edge in the graph used to predict that edge.
  • Do not count duplicated neighbor IDs as independent evidence.
  • Do not sample more hops than the model can use.

Read next

ai-data
graph-machine-learning
Storage details