Skip to content
AITroveRead. Build. Understand.
Make this comfortable

References to groups and related entities: identity versus bridging

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

“They” may denote two earlier actors, while “the gateway” may be associated with a deployment without being that deployment. Model both cases explicitly.

Separate identity from association

Coreference links mentions that denote the same entity. Bridging links a new entity to something already mentioned: “the deployment failed; the gateway retried” does not mean gateway and deployment are identical. Record a typed relation instead of merging their chains. A bad merge can make a later summary claim the deployment itself retried. Mention-chain identity provides the base rule; bridging preserves a related but distinct referent.

Allow several antecedents when the text does

“The API and worker restarted. They recovered after 47 seconds” refers to a group formed from two earlier entities. Picking the nearest single noun loses one participant. Represent the group mention with a stable group ID and member IDs, while keeping both original entity chains intact. Do not invent group membership from adjacency alone. The plural form, sentence structure and incident context are evidence, but a reviewer should inspect high-impact recovery claims.

Preserve unresolved references

A note that says “it failed again” may follow mentions of a queue, a gateway and a check. If the context does not identify which one failed, store an unresolved mention with candidate chain IDs and its source span. Do not attach it to the most recent noun just to complete the graph. Role frames can help identify what action a mention performs, but they cannot recover an absent antecedent from thin air.

Score the distinctions that matter

Evaluate identity-chain links, group membership, bridging edges and unresolved rate separately. An ordinary pair score can reward a system that overmerges easy names while missing group and bridging cases. Include repeated component names, revisions and quoted prior updates. Inspect whether downstream timelines and summaries attribute failures to the right service. The reference audit project checks the user-visible effect rather than only mention-pair agreement.

Implementation

python
def reference_record(mention_id, relation, targets, source_revision):
    if not mention_id or not source_revision:
        raise ValueError("mention and source revision are required")
    if relation == "group" and len(set(targets)) < 2:
        raise ValueError("a group needs distinct members")
    if relation == "identity" and len(targets) != 1:
        raise ValueError("identity needs one chain")
    if relation not in {"group", "identity", "bridging", "unresolved"}:
        raise ValueError("unknown reference relation")
    return {"mention_id": mention_id, "relation": relation,
            "targets": tuple(targets), "source_revision": source_revision}

group = reference_record("they-47", "group", ["api-12", "worker-19"], "note-r3")
assert group["targets"] == ("api-12", "worker-19")
assert reference_record("it-82", "unresolved", [], "note-r3")["relation"] == "unresolved"

Performance and operating cost

Validating k target IDs costs O(k) time and O(k) temporary space. Candidate generation over m mentions can reach O(m²) if every mention is compared with every earlier one; section and incident boundaries reduce that work. Human review is most valuable where a group or bridging decision changes a published incident claim.

Common Mistakes

  • Merging a component with an event merely because they are related.
  • Forcing a plural pronoun onto one nearest antecedent.
  • Replacing an unresolved reference with a guess to make a graph complete.
  • Reporting only overall chain scores while group membership is wrong.

Read next

ai-data
natural-language-processing
Storage details