Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Coreference chains: track mentions without guessing identity

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A document may call the same service “the gateway,” “it” and “that component.” Link mentions within the document before making cross-document identity claims.

Mention identity is local

A mention is a span in a specific source revision. A coreference chain groups mentions that refer to one discourse entity in that source. It is not a global database identity. In “The gateway retried the charge. It timed out,” “It” may mean the gateway or the charge operation; a model should preserve uncertainty when context does not resolve it. Record span offsets, chain ID, confidence and adjudication version. Span-offset annotation provides the source coordinate system.

Separate coreference from entity linking

Coreference answers whether two mentions in this text refer together. Entity linking asks whether a mention maps to a known catalog record. An internal service called “Atlas” may share its name with several products; do not merge chains across incidents by surface string. First build document-local chains, then link a chain to a catalog ID using context and a candidate list. Keep NIL as an explicit result when none fits. A wrong global link can poison later relation extraction more than an unresolved pronoun does.

Annotate difficult discourse

Review pronouns, shortened names, nominal phrases, plural references and quoted text. Mark genuinely ambiguous cases separately rather than forcing agreement. Split evaluation by source document or incident; copied paragraphs in both train and test will overstate generalization. Score mention detection and chain linking separately. A system can find every pronoun yet attach them to the wrong named entity, so inspect chain-level errors with the original context.

Keep provenance through changes

A corrected document revision can shift every offset. Store mention records with revision IDs and recompute spans only after a verified mapping. Never let chain IDs silently survive a rewrite. In production, show the source sentence and antecedent when a resolved entity drives a downstream action. The relation ledger project uses that record to avoid turning a weak pronoun guess into a firm operational fact.

Implementation

python
def validate_mentions(source_text, mentions):
    chain_members = {}
    for mention in mentions:
        start, end = mention["start"], mention["end"]
        if not 0 <= start < end <= len(source_text):
            raise ValueError("mention is outside the source revision")
        if source_text[start:end] != mention["surface"]:
            raise ValueError("mention no longer matches the source")
        chain_members.setdefault(mention["chain_id"], []).append((start, end))
    return chain_members

incident_note = "The gateway retried. It failed."
mentions = [{"start": 4, "end": 11, "surface": "gateway", "chain_id": "c47"},
            {"start": 21, "end": 23, "surface": "It", "chain_id": "c47"}]
assert len(validate_mentions(incident_note, mentions)["c47"]) == 2

Performance and operating cost

Validation is O(m + selected characters) time and O(m) space for m mentions. Candidate antecedent search can approach O(m²) if every mention is compared with every earlier mention; prune by sentence distance and compatible type only after measuring recall on long documents. Global entity linking adds catalog retrieval cost. Store NIL and ambiguity rates because faster forced linking can increase expensive downstream corrections.

Common Mistakes

  • Treating a shared name as proof of global identity.
  • Counting mention detection as successful coreference.
  • Reusing offsets after the source revision changes.
  • Forcing one antecedent where the document leaves the reference ambiguous.

Read next

Continue the workflow: Relation extraction with direction, negation and evidence.

Continue the workflow: Dialogue state: apply slot updates, corrections and deletions.

Continue the workflow: References to groups and related entities: identity versus bridging.

ai-data
natural-language-processing
Storage details