Similarity is not entailment. A claim about a completed refund needs source evidence that supports completion, not merely a passage containing the same order ID.
Textual entailment: premise, hypothesis and evidence boundaries
Name the two texts
The premise is a versioned source passage; the hypothesis is the claim being checked. Entailment means the premise supports the hypothesis under the task policy, contradiction means it supports an incompatible statement, and neutral means neither follows. A sentence about “refund requested” is neutral to “refund completed,” even if the phrases are close in an embedding space. Preserve document revision, speaker and time with each pair. Summary evidence depends on this distinction.
Define the scope of inference
A premise can be insufficient because the relevant fact lives in another paragraph, or because the source itself is ambiguous. Keep a separate retrieval-evidence label so a neutral model output is not automatically interpreted as “the fact is false.” Include negation, quantities, temporal qualifiers, ownership and modal verbs in the audit set. Annotate the minimal source span that justifies a non-neutral label. If reviewers cannot agree, mark the case disputed rather than forcing a training class.
Avoid shortcut labels
Models can exploit lexical overlap and domain templates. Build counterexamples where the same identifiers appear with opposite status or where the relation direction reverses. Split by incident, document family and time. If a template sentence appears in both train and final audit, apparent performance may come from memorizing its wording. Grouped temporal validation applies to premise-hypothesis pairs just as it does to ticket classifiers.
Gate operational use
Report per-class precision and recall, unsupported-claim false negatives, abstention coverage and error slices. An entailment score is a screening signal, not automatic proof. Require the source span to round-trip, the revision to be current and access to be authorized. For high-impact claims, route ambiguous scores to a reviewer. Conflict review adds corrections and time; the project tests the full chain.
Implementation
def inference_record(premise_id, premise_revision, hypothesis, label, evidence_span):
if label not in {"entails", "contradicts", "neutral"}:
raise ValueError("unknown inference label")
if not premise_id or not premise_revision or not hypothesis.strip():
raise ValueError("source and claim are required")
if label != "neutral" and evidence_span is None:
raise ValueError("non-neutral label needs source evidence")
return {"premise_id": premise_id, "revision": premise_revision,
"hypothesis": hypothesis, "label": label, "evidence_span": evidence_span}
claim = inference_record("incident-47", "note-v3", "The retry completed.",
"entails", (12, 27))
assert claim["label"] == "entails"
Performance and operating cost
The record validator is O(1) time and space apart from copying the claim. Checking many hypotheses against many passages can be O(c·p) model calls for c claims and p passages; retrieval should narrow candidates, but report its recall before claiming verifier quality. Human adjudication of ambiguous cases is the main editorial cost. Keep evidence and access checks even if they add serving latency.
Common Mistakes
- Treating semantically similar text as proof of a claim.
- Interpreting neutral as a contradiction or as a false fact.
- Letting a non-neutral label through without a source span.
- Evaluating only clean pairs while ignoring retrieval and revision failures.
Read next
- Review claim conflicts across time and source revisions
- Project: verify incident claims against reviewed source revisions
- Document summaries with sentence-level evidence contracts
- QA answerability: calibrate abstention and evidence quality
- Text validation: split conversations, duplicates and time together
Continue the workflow: Review claim conflicts across time and source revisions.
Continue the workflow: Paraphrase equivalence: make hard negatives change the decision.
Continue the workflow: Generated answers: claim evidence and unanswered state.
Continue the workflow: Text simplification evaluation: comprehension and release slices.
