Extract what a document argues as typed spans and relations, rather than treating every confident sentence as supported evidence.
Argument units: claims, premises and evidence spans
Choose an annotation unit
A change proposal can state “Extend the timeout to 82 minutes,” then cite an incident count as its reason. The recommendation is a claim; the count is a potential premise or evidence span. Record offsets, speaker, document revision and unit type. A sentence can contain more than one unit, while a premise can span sentences. Keep the exact source text so a reviewer can challenge a model’s extraction. Span contracts prevent normalized text from drifting away from the document that was annotated.
Separate roles from truth
Calling a sentence evidence means that it is offered to support or attack a claim; it does not certify that the statistic is correct. Store the evidence source, measurement window and verification state separately. Likewise, a claim may be quoted by an author who rejects it. Mark attribution and quotation boundaries before assigning the author a stance. Attribution review guards against converting a quoted objection into the proposal owner’s view.
Build typed relations
Connect premise and claim IDs with support, attack or unresolved links. Do not infer a support edge from proximity alone: “82 incidents occurred” may be a background fact, not a reason for raising a timeout. A relation can cross paragraph boundaries and may depend on a missing assumption. Preserve reviewer rationale and source revision for each edge. NLI evidence tests entailment, while argument relations represent how a writer uses one unit in an argument; those are related but not identical decisions.
Audit missing structure
Evaluate span boundaries, unit type, relation direction, attribution and evidence verification separately. Include proposals that give no reason, contain competing reasons or cite a disproven metric. A graph with every sentence connected can look complete while inventing support. The target stance lesson adds the position each speaker takes; the project checks whether a release summary respects the map.
Implementation
def argument_unit(document_id, revision_id, source_text, start, end,
unit_type, speaker_id):
allowed = {"claim", "premise", "evidence"}
if unit_type not in allowed or not (0 <= start < end <= len(source_text)):
raise ValueError("invalid argument span or type")
if not document_id or not revision_id or not speaker_id:
raise ValueError("source identity and speaker are required")
return {"document_id": document_id, "revision_id": revision_id,
"start": start, "end": end, "text": source_text[start:end],
"unit_type": unit_type, "speaker_id": speaker_id}
proposal = "Extend the gateway-west timeout to 82 minutes."
unit = argument_unit("change-47", "r3", proposal, 0, len(proposal),
"claim", "operator-82")
assert unit["text"] == proposal
assert unit["unit_type"] == "claim"
Performance and operating cost
Validating the span bounds is O(1); slicing a unit of s characters takes O(s) time and space for the copied text. A real annotation store also needs versioned offsets and relation indexing. This example records an offered role, not factual verification of the unit or automatic discovery of its boundaries.
Common Mistakes
- Calling an offered statistic verified evidence without checking its source.
- Assuming one sentence always equals one argument unit.
- Assigning a quoted objection to the quoting author.
- Drawing a support edge from sentence proximity alone.
