A fluent summary can insert a false date or owner. Require every operational claim to point back to a source passage and mark unsupported claims for review.
Document summaries with sentence-level evidence contracts
Choose the summary job
A handoff summary for an incident differs from an abstract for a report. Define audience, maximum length, required fields, prohibited inferences and the source snapshot. For a support handoff, a good contract might require customer impact, confirmed cause, current mitigation and unresolved questions. If a field is absent in source material, write “not recorded” instead of filling it with a likely value. The choice of extractive or generated text follows from this contract, not from a preference for a model family.
Attach evidence to claims
Split a candidate summary into claims or sentences and associate each with one or more source spans. A claim about a time, count, person or cause needs exact supporting text from the correct revision. A citation marker alone is insufficient if the passage does not entail the statement. Keep source IDs, revision IDs and offsets in the internal representation; render a reader-friendly reference only after validation. Span offsets provide the mechanical link, and entity evaluation highlights where a single character error can change meaning.
Separate coverage from faithfulness
A summary can be perfectly faithful yet omit the only action item. Conversely, it can cover every topic while inventing a resolution. Judge both: did the summary include required facts, and are its stated facts supported? Surface contradictions and temporal order separately. Automated similarity metrics are useful screens but cannot alone establish factual support, especially for short numeric changes. Review a stratified sample of summaries and hard cases with a rubric before release.
Handle long sources
For a long incident log, select relevant windows with timestamps and speaker identity before drafting. Ensure the selected windows include later corrections: an early hypothesis may be explicitly disproved. Preserve the source order so “mitigation attempted” cannot become “mitigation succeeded.” If no window contains enough evidence, abstain or request a human summary. The claim verification path turns that decision into a check, and the project uses it in a release workflow.
Implementation
def validate_evidence_spans(source_text, claims):
checked = []
for claim in claims:
supports = claim.get("support_spans", ())
if not supports:
raise ValueError("summary claim lacks source evidence")
for start, end, expected in supports:
if not 0 <= start < end <= len(source_text):
raise ValueError("source span is out of range")
if source_text[start:end] != expected:
raise ValueError("source evidence changed")
checked.append(claim["text"])
return checked
source = "Callback retry completed at 14:47."
assert validate_evidence_spans(source, [{"text": "The retry completed.",
"support_spans": [(9, 24, "retry completed")]}])
Performance and operating cost
Span validation takes O(c + total evidence characters) time and O(c) output space for c claims. Model generation is often the dominant cost, especially when long source windows are passed repeatedly. Select windows once, cap input length, and measure how often selection omits a later correction. This mechanical validator confirms that cited text exists; it cannot prove that the claim follows from that text, so human review remains a release gate for consequential summaries.
Common Mistakes
- Treating fluent wording as evidence of correctness.
- Accepting any source span as support without checking the claim’s meaning.
- Summarizing an early hypothesis after a later correction disproved it.
- Measuring only overlap with a reference summary and missing unsupported claims.
Read next
- Verify summary claims, corrections and human-review triggers
- Project: publish evidence-backed incident handoff summaries
- Entity spans: align annotations to the original text
- Decode entities and evaluate exact spans, not token accuracy
- Text inference: package tokenizer, labels and reject paths
Continue the workflow: Verify summary claims, corrections and human-review triggers.
Continue the workflow: QA answerability: calibrate abstention and evidence quality.
Continue the workflow: Textual entailment: premise, hypothesis and evidence boundaries.
