A document may call the same service “the gateway,” “it” and “that component.” Link mentions within the document before making cross-document identity claims.
Coreference chains: track mentions without guessing identity
Mention identity is local
A mention is a span in a specific source revision. A coreference chain groups mentions that refer to one discourse entity in that source. It is not a global database identity. In “The gateway retried the charge. It timed out,” “It” may mean the gateway or the charge operation; a model should preserve uncertainty when context does not resolve it. Record span offsets, chain ID, confidence and adjudication version. Span-offset annotation provides the source coordinate system.
Separate coreference from entity linking
Coreference answers whether two mentions in this text refer together. Entity linking asks whether a mention maps to a known catalog record. An internal service called “Atlas” may share its name with several products; do not merge chains across incidents by surface string. First build document-local chains, then link a chain to a catalog ID using context and a candidate list. Keep NIL as an explicit result when none fits. A wrong global link can poison later relation extraction more than an unresolved pronoun does.
Annotate difficult discourse
Review pronouns, shortened names, nominal phrases, plural references and quoted text. Mark genuinely ambiguous cases separately rather than forcing agreement. Split evaluation by source document or incident; copied paragraphs in both train and test will overstate generalization. Score mention detection and chain linking separately. A system can find every pronoun yet attach them to the wrong named entity, so inspect chain-level errors with the original context.
Keep provenance through changes
A corrected document revision can shift every offset. Store mention records with revision IDs and recompute spans only after a verified mapping. Never let chain IDs silently survive a rewrite. In production, show the source sentence and antecedent when a resolved entity drives a downstream action. The relation ledger project uses that record to avoid turning a weak pronoun guess into a firm operational fact.
Implementation
def validate_mentions(source_text, mentions):
chain_members = {}
for mention in mentions:
start, end = mention["start"], mention["end"]
if not 0 <= start < end <= len(source_text):
raise ValueError("mention is outside the source revision")
if source_text[start:end] != mention["surface"]:
raise ValueError("mention no longer matches the source")
chain_members.setdefault(mention["chain_id"], []).append((start, end))
return chain_members
incident_note = "The gateway retried. It failed."
mentions = [{"start": 4, "end": 11, "surface": "gateway", "chain_id": "c47"},
{"start": 21, "end": 23, "surface": "It", "chain_id": "c47"}]
assert len(validate_mentions(incident_note, mentions)["c47"]) == 2
Performance and operating cost
Validation is O(m + selected characters) time and O(m) space for m mentions. Candidate antecedent search can approach O(m²) if every mention is compared with every earlier mention; prune by sentence distance and compatible type only after measuring recall on long documents. Global entity linking adds catalog retrieval cost. Store NIL and ambiguity rates because faster forced linking can increase expensive downstream corrections.
Common Mistakes
- Treating a shared name as proof of global identity.
- Counting mention detection as successful coreference.
- Reusing offsets after the source revision changes.
- Forcing one antecedent where the document leaves the reference ambiguous.
Read next
- Entity spans: align annotations to the original text
- Relation extraction with direction, negation and evidence
- Project: build a reviewed entity-relation ledger from incident notes
- Decode entities and evaluate exact spans, not token accuracy
- Knowledge Graphs Tutorial
Continue the workflow: Relation extraction with direction, negation and evidence.
Continue the workflow: Dialogue state: apply slot updates, corrections and deletions.
Continue the workflow: References to groups and related entities: identity versus bridging.
