A comma change and a new rollback threshold should not trigger the same review. Compare claim fields and their evidence across source revisions.
Semantic document diffs: changed facts versus wording edits
Extract claims before judging significance
A runbook revision changes “wait 47 minutes” to “wait 82 minutes.” The sentence structure barely changes, but the operational instruction does. Represent each claim with subject, action, condition, quantity, unit, source span and revision. A whitespace or grammar edit may leave those fields unchanged. Do not rank edit impact solely by character distance. Quantity spans make the changed threshold visible as data rather than as a token difference.
Keep old and new evidence side by side
Align claims by stable section and entity identity, then compare their fields. A deleted step, added exception or changed negation can alter behavior even if most words stay the same. Keep both source passages and mark the proposed change type: surface, factual, policy or unresolved. An automated diff is a review aid, not an authority to declare equivalence. Predicate and negation scope matter when “disable retries” becomes “do not disable retries.”
Track applicability and supersession
A revised production step may leave the staging step untouched. Record environment, effective date and explicit supersession so the impact analysis does not invalidate unrelated procedures. If the new text is ambiguous, preserve the previous active claim until the owner resolves release policy; never blend both into one instruction. Conflict handling can expose two active claims that now disagree.
Evaluate operational consequences
Review threshold changes, negation flips, added exceptions, removed prerequisites and wording-only edits. Score changed-claim detection and false alarms separately. Then test whether search answers and summaries stop using an invalidated instruction. The revision project audits derived artifacts after an edit, not just the text diff shown to an editor.
Implementation
def classify_claim_edit(before, after):
if before["claim_id"] != after["claim_id"]:
return "review-alignment"
policy_fields = {"action", "condition", "quantity", "unit", "negated"}
changed = {field for field in policy_fields if before[field] != after[field]}
if changed:
return "meaning-change"
return "surface-or-context-change" if before["text"] != after["text"] else "unchanged"
old = {"claim_id": "rollback-wait", "action": "wait", "condition": "after rollback",
"quantity": 47, "unit": "min", "negated": False, "text": "Wait 47 minutes."}
new = {**old, "quantity": 82, "text": "Wait 82 minutes."}
assert classify_claim_edit(old, new) == "meaning-change"
assert classify_claim_edit(old, {**old, "text": "Wait for 47 minutes."}) == "surface-or-context-change"
Performance and operating cost
Comparing f structured fields is O(f) time and O(f) space for the changed-field set. Extracting and aligning claims from whole documents is the expensive step, and a stable ID can itself be wrong after a section rewrite. Treat “surface-or-context-change” as needing contextual review, not as automatic semantic equivalence.
Common Mistakes
- Calling a numeric threshold edit minor because few characters changed.
- Ignoring a negation flip in an otherwise similar sentence.
- Applying a production change to staging evidence.
- Treating a heuristic surface label as proof that meaning is unchanged.
