Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Entity linking: candidate generation and the NIL decision

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A detected name is only a span. Linking it to a product or service record requires context, scope and an option for no match.

Separate detection from identity

A mention such as “Atlas” could be a service, a dashboard, a customer project or an ordinary word. Entity recognition supplies the span and type; linking selects a catalog ID, or NIL when no record fits. Keep the original mention, sentence context, source revision and tenant scope. A string match is a candidate generator, not an identity proof. Entity spans define the input boundary.

Build candidates without collapsing aliases

Index approved names, aliases and retired names by catalog version. Retrieve candidates using exact, normalized and context terms, while preserving case-sensitive codes. Constrain by entity type and authorized tenant before ranking. Do not let a common alias from another customer account leak through a suggestion list. A new service that has not yet been registered needs NIL, rather than the closest old record. Coreference can gather nearby descriptions but does not create a catalog ID.

Compare context and reject uncertainty

Score candidate type, neighboring words, valid deployment period and distinct identifiers. A mention in an incident about the billing gateway should not link to a storage dashboard named the same way. Calibrate a reject threshold on reviewed mentions, including unseen entities. If the top candidates are close, send them for review with concise evidence. Store the chosen catalog revision so a later rename does not silently change the original interpretation.

Evaluate the full path

Measure detection recall, candidate recall, correct-ID precision, NIL precision and wrong cross-tenant links separately. Group related incident documents before splitting. A ranker cannot recover the right entity if candidate generation omitted it. Catalog revisions address merges and deletions; the project tests the whole service.

Implementation

python
def choose_entity_link(mention, candidates, minimum_score, margin):
    visible = [candidate for candidate in candidates
               if candidate["tenant_id"] == mention["tenant_id"]]
    ranked = sorted(visible, key=lambda candidate: candidate["score"], reverse=True)
    if not ranked or ranked[0]["score"] < minimum_score:
        return None, "nil"
    if len(ranked) > 1 and ranked[0]["score"] - ranked[1]["score"] < margin:
        return None, "review-ambiguity"
    return ranked[0]["entity_id"], "linked"

mention = {"text": "Atlas", "tenant_id": "north"}
options = [{"entity_id": "svc-47", "tenant_id": "north", "score": 0.91},
           {"entity_id": "svc-82", "tenant_id": "south", "score": 0.98}]
assert choose_entity_link(mention, options, 0.8, 0.08) == ("svc-47", "linked")

Performance and operating cost

Filtering c candidates is O(c), and sorting takes O(c log c) time with O(c) space. A top-two selection would be O(c) if full ranking is unnecessary. Candidate retrieval and human review dominate system cost. Measure missed correct candidates and wrongful links independently; a low-latency ranker cannot repair a missing catalog entry.

Common Mistakes

  • Equating a recognized mention with a verified entity ID.
  • Forcing an unseen product to match the nearest known one.
  • Showing aliases from another tenant in candidate suggestions.
  • Evaluating rank accuracy only on mentions whose correct ID was retrieved.

Read next

Continue the workflow: Domain word senses: inventory, abstention and new meanings.

Continue the workflow: Cross-document reference chains: revisions, scope and access.

Continue the workflow: Natural-language queries: intent, schema grounding and ambiguity.

Continue the workflow: Transliteration variants are candidates, not identity.

ai-data
natural-language-processing
Storage details