Graph structure is a modeled data contract; adjacency choices affect both the learned signal and the cost of traversal.
Graph construction: direction, duplicates and high-degree nodes
Represent relation meaning
A prerequisite edge from lesson A to lesson B can mean A must precede B, while a learner-to-lesson completion edge has a different direction and time. Do not symmetrize directed relations merely because a library expects undirected input. If reverse edges are added for message passing, give them a separate relation type so the original meaning stays visible.
Remove false multiplicity
Repeated page views do not necessarily create new learner-lesson relationships. Decide whether an edge represents existence, count or weighted recency. A graph with duplicate edges can amplify one learner’s activity in aggregation. Snapshot identity should make retries and genuine repeated behavior distinguishable.
Inspect hubs
A popular concept node may connect to thousands of lessons. Naive neighborhood expansion from it can dominate memory and drown out rare prerequisite paths. Report degree distribution and isolate suspicious supernodes such as “all lessons.” Neighbor sampling or relation-specific caps control cost, but they can hide useful low-volume links; evaluate that loss.
Test a tiny graph
Create three lesson nodes, one concept node and four directed edges. Add a duplicate edge and verify unique degree does not change. Reverse one prerequisite and confirm path traversal changes, while a completion edge remains in its own relation. A graph export should preserve isolated lessons as nodes, even when they have no interactions yet.
Implementation
def unique_out_neighbors(edge_rows, source_id, relation):
return {edge["target"] for edge in edge_rows
if edge["source"] == source_id and edge["relation"] == relation}Performance and operating cost
A scan of E edges costs O(E) time and O(D) memory for D unique matching neighbors. An adjacency index can make repeated lookups closer to O(D), at the cost of O(E) stored edges and careful update handling.
Common Mistakes
- Do not reverse a prerequisite relation without labeling the reverse direction.
- Do not let duplicate events inflate degree or message weight silently.
- Do not report only average degree when a few hubs dominate work.
Read next
- Graph learning contract: entities, relations and snapshot time
- Neighborhood features: aggregate only the graph that existed at prediction time
- Graph serving: handle new nodes, edge deletion and snapshot swaps
- Content and collaborative signals: give sparse learners a real fallback
Continue the workflow: Graph validation: type constraints, cardinality and quality gates.
