Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Entity resolution: aliases, merge decisions and reversible identity

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Entity resolution joins records referring to the same real entity, but a false merge can corrupt every downstream path.

Separate candidates from decisions

Two lesson pages with similar names may describe different topics. Generate candidate matches from normalized titles, aliases and subject context, then require stronger evidence such as canonical path history or shared stable ID to merge. Store match score, rule version and reviewer decision. The identity contract must say whether a split or merge changes public paths and historical claims.

Keep aliases typed

An alias may be a former title, abbreviation, spelling variant or disambiguation label. Do not let an alias become a unique key without checking collisions. A concept named “Java streams” can refer to a language API, while another document uses “stream” for event processing. Subject and relation context can prevent a false match.

Make merges reversible

Preserve source entity IDs and redirects when merging. If a reviewer later finds the merge wrong, claims must be separable by original evidence rather than copied irreversibly to the surviving node. Track the time each identity decision became active so a historical graph snapshot can be reproduced. Replay and correction rules apply to identity events too.

Test the collision

Create lessons “Streaming in Java” and “Streaming Analytics,” both aliased as “streams.” A title-only resolver would merge them. Require subject and canonical path context, then keep them separate. Add a genuine renamed lesson with two old paths and confirm it resolves to one entity while retaining the old identifiers as audit links.

Implementation

python
def alias_candidates(alias_rows, query_alias, subject):
    normalized = query_alias.casefold().strip()
    return [row["entity_id"] for row in alias_rows
            if row["alias"].casefold().strip() == normalized
            and row["subject"] == subject]

Performance and operating cost

A scan of A aliases costs O(A) time and O(M) output space for M candidates. Indexing normalized alias plus subject makes lookup faster, but collision review and reversible merge history remain essential.

Common Mistakes

  • Do not merge entities on title similarity alone.
  • Do not assume an alias is globally unique.
  • Do not discard source IDs and evidence during a merge.

Read next

ai-data
knowledge-graphs
Storage details