Entity resolution joins records referring to the same real entity, but a false merge can corrupt every downstream path.
Entity resolution: aliases, merge decisions and reversible identity
Separate candidates from decisions
Two lesson pages with similar names may describe different topics. Generate candidate matches from normalized titles, aliases and subject context, then require stronger evidence such as canonical path history or shared stable ID to merge. Store match score, rule version and reviewer decision. The identity contract must say whether a split or merge changes public paths and historical claims.
Keep aliases typed
An alias may be a former title, abbreviation, spelling variant or disambiguation label. Do not let an alias become a unique key without checking collisions. A concept named “Java streams” can refer to a language API, while another document uses “stream” for event processing. Subject and relation context can prevent a false match.
Make merges reversible
Preserve source entity IDs and redirects when merging. If a reviewer later finds the merge wrong, claims must be separable by original evidence rather than copied irreversibly to the surviving node. Track the time each identity decision became active so a historical graph snapshot can be reproduced. Replay and correction rules apply to identity events too.
Test the collision
Create lessons “Streaming in Java” and “Streaming Analytics,” both aliased as “streams.” A title-only resolver would merge them. Require subject and canonical path context, then keep them separate. Add a genuine renamed lesson with two old paths and confirm it resolves to one entity while retaining the old identifiers as audit links.
Implementation
def alias_candidates(alias_rows, query_alias, subject):
normalized = query_alias.casefold().strip()
return [row["entity_id"] for row in alias_rows
if row["alias"].casefold().strip() == normalized
and row["subject"] == subject]Performance and operating cost
A scan of A aliases costs O(A) time and O(M) output space for M candidates. Indexing normalized alias plus subject makes lookup faster, but collision review and reversible merge history remain essential.
Common Mistakes
- Do not merge entities on title similarity alone.
- Do not assume an alias is globally unique.
- Do not discard source IDs and evidence during a merge.
