Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Discourse units: relation labels and evidence scope

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A relation such as contrast or cause connects claims, not arbitrary adjacent sentences. Preserve the two spans and the source of the link.

Segment before naming a relation

“The queue drained, but retries remained high” contains two claims and an explicit contrast. Store each argument span, its document revision and the connective span. A sentence can contain several units; a paragraph can also support a relation across sentence boundaries. Do not assume punctuation determines the boundary. Document section boundaries prevent an incident note from being connected to the next unrelated case.

Keep relation types distinct

Contrast says the claims differ in an expected dimension. Cause says one event produced another. Temporal order only says one happened before another. Elaboration adds detail without changing the underlying claim. Attribution records who made a statement, including uncertainty. One pair may carry more than one plausible relation; preserve separate proposals and evidence rather than forcing a single exclusive label. Event-order logic can validate temporal edges without pretending they are causal.

Track the evidence direction

“Rollback followed the alarm” states order, not proof that the alarm caused rollback. For an explicit causal connective, record which argument is the stated cause and which is the effect; the arrow matters. If the source merely quotes an operator’s hypothesis, attribution must remain attached to the edge. A model-generated causal relation cannot turn a quoted suspicion into an established incident fact. Relation provenance stores its supporting span and revision.

Evaluate spans and relations separately

Measure unit boundaries, relation class, direction, attribution and unsupported causal edges on reviewed documents. A system may select the right label while attaching it to the wrong clause. Include multi-clause status updates, corrections and documents with repeated service names. Reviewers should see source text beside every proposed edge. The discourse-map project uses this distinction in an incident handoff.

Implementation

python
def relation_edge(left_span, right_span, label, evidence_span, revision):
    if not left_span or not right_span or not evidence_span or not revision:
        raise ValueError("both arguments, evidence and revision are required")
    if label not in {"contrast", "cause", "temporal", "elaboration", "attribution"}:
        raise ValueError("unknown relation label")
    return {"left": left_span, "right": right_span, "label": label,
            "evidence": evidence_span, "source_revision": revision}

edge = relation_edge("queue drained", "retries remained high", "contrast",
                     "but", "update-r47")
assert edge["label"] == "contrast"

Performance and operating cost

Constructing one edge is O(1); segmenting n tokens is at least O(n). Naively comparing all u units produces O(u²) candidate pairs, so restrict candidates with document and section boundaries before classifying. Human review cost rises sharply for causal edges because a wrong causal claim can distort remediation. Track unsupported-cause rate separately from ordinary relation accuracy.

Common Mistakes

  • Treating “after” as proof of causality.
  • Joining claims from different incident sections.
  • Dropping the speaker on a quoted hypothesis.
  • Reporting relation accuracy without checking argument boundaries.

Read next

Continue the workflow: Abusive-language labels: context, target and quotation.

Continue the workflow: Operational rules: actors, triggers, obligations and exceptions.

ai-data
natural-language-processing
Storage details