Extract service dependencies from incident notes while preserving mention chains, evidence revisions and a human path for disputed edges.
Project: build a reviewed entity-relation ledger from incident notes
The ledger boundary
The product team wants to ask which services depended on a queue during an incident. Do not insert model output directly into the authoritative service catalog. Create a ledger of proposed assertions with source document revision, entity mentions, relation direction, effective time and review state. A “not stated” result is valid. Restrict the relation vocabulary to operational questions the reviewer can actually verify.
Prepare reviewed documents
Capture incident notes before any cleanup that shifts character offsets. Annotate service and queue spans, local mention chains, candidate catalog IDs and relation evidence. Include negated and superseded statements. Group notes from the same incident in one split and reserve later incidents for final audit. Reviewers should settle whether “it” refers to the service or the queue before labeling the edge. Follow mention identity and relation provenance in that order.
Gate the extraction
Evaluate every stage, including exact entity spans, chain links, catalog links, candidate pair coverage and end-to-end reviewed edges. Set separate thresholds for false dependencies and missed dependencies; they have different costs. Send ambiguous pronouns, unknown catalog identities and negated assertions to review. Compare a simple rule baseline against the trained model. A gain in classifier F1 does not justify release if candidate generation loses rare cross-sentence edges.
Operate revisions
A reviewer can approve, reject or supersede an assertion. On a document revision, new proposed edges point to the new source digest; old approved edges remain in history but may be retired from the current graph view. Audit queries must filter state and time. Keep raw incident note permissions attached to evidence snippets. The published ledger exposes structured facts only to users authorized for the underlying documents.
Implementation
def current_approved_edges(assertions, visible_source_revisions):
return [assertion for assertion in assertions
if assertion["state"] == "approved"
and assertion["source_revision"] in visible_source_revisions
and assertion.get("superseded_by") is None]
ledger = [
{"edge_id": "edge-47", "state": "approved", "source_revision": "note-v7", "superseded_by": None},
{"edge_id": "edge-48", "state": "staged", "source_revision": "note-v8", "superseded_by": None},
]
assert [row["edge_id"] for row in current_approved_edges(ledger, {"note-v7"})] == ["edge-47"]
Performance and operating cost
The in-memory view scan is O(n) time for n assertions and O(v) output space for v visible current edges. A database should index review state, source revision and supersession for larger ledgers. Extraction candidate generation and human adjudication dominate production cost. Measure review queue age and false approved edges, not only model throughput, before promoting extracted relations into a graph view.
Common Mistakes
- Writing predicted edges directly into the authoritative catalog.
- Dropping old edge history when a note is corrected.
- Showing evidence text to users who lack source access.
- Evaluating only the relation classifier while ignoring coreference and entity linking errors.
Read next
- Coreference chains: track mentions without guessing identity
- Relation extraction with direction, negation and evidence
- Project: ship an auditable support-entity extractor
- Knowledge Graphs Tutorial
- Text inference: package tokenizer, labels and reject paths
Continue the workflow: Project: extract incident actions from validated syntax.
Continue the workflow: Entity links across catalog renames, merges and deletions.
