Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Graph construction: direction, duplicates and high-degree nodes

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Graph structure is a modeled data contract; adjacency choices affect both the learned signal and the cost of traversal.

Represent relation meaning

A prerequisite edge from lesson A to lesson B can mean A must precede B, while a learner-to-lesson completion edge has a different direction and time. Do not symmetrize directed relations merely because a library expects undirected input. If reverse edges are added for message passing, give them a separate relation type so the original meaning stays visible.

Remove false multiplicity

Repeated page views do not necessarily create new learner-lesson relationships. Decide whether an edge represents existence, count or weighted recency. A graph with duplicate edges can amplify one learner’s activity in aggregation. Snapshot identity should make retries and genuine repeated behavior distinguishable.

Inspect hubs

A popular concept node may connect to thousands of lessons. Naive neighborhood expansion from it can dominate memory and drown out rare prerequisite paths. Report degree distribution and isolate suspicious supernodes such as “all lessons.” Neighbor sampling or relation-specific caps control cost, but they can hide useful low-volume links; evaluate that loss.

Test a tiny graph

Create three lesson nodes, one concept node and four directed edges. Add a duplicate edge and verify unique degree does not change. Reverse one prerequisite and confirm path traversal changes, while a completion edge remains in its own relation. A graph export should preserve isolated lessons as nodes, even when they have no interactions yet.

Implementation

python
def unique_out_neighbors(edge_rows, source_id, relation):
    return {edge["target"] for edge in edge_rows
            if edge["source"] == source_id and edge["relation"] == relation}

Performance and operating cost

A scan of E edges costs O(E) time and O(D) memory for D unique matching neighbors. An adjacency index can make repeated lookups closer to O(D), at the cost of O(E) stored edges and careful update handling.

Common Mistakes

  • Do not reverse a prerequisite relation without labeling the reverse direction.
  • Do not let duplicate events inflate degree or message weight silently.
  • Do not report only average degree when a few hubs dominate work.

Read next

Continue the workflow: Graph validation: type constraints, cardinality and quality gates.

ai-data
graph-machine-learning
Storage details