Create reviewed negation, threshold and paraphrase pairs for operational answers, then hold a release when a changed instruction receives the old decision.
Project: gate a runbook answer release with contrast cases
Define the operational risk
A runbook answer system recommends rollback actions. An old passage permits rollback after 47 minutes; a revision requires 82 minutes and adds a production-only block. Build cases that distinguish the versions and environments. Include a harmless paraphrase that should preserve the answer. The goal is not merely to raise average answer accuracy: it is to catch a system that repeats an obsolete permission when a small policy change matters. Semantic revision review establishes which claims changed.
Freeze the contrast set
Store base and edited passages, user question, environment, source revision, expected decision and reviewer rationale. Include negation flips, threshold crossings, unchanged paraphrases and distractor text from another service. Keep all variants of an incident in one evaluation family. Review the expected labels without seeing model predictions. Edit contracts prevent a test from quietly changing several variables at once.
Exercise the full answer path
Run retrieval and answer generation against the correct active runbook revision. Capture retrieved passage IDs, generated decision, confidence or abstention state and any citation spans used internally for verification. A model may fail because it retrieves a stale index entry, misses the negation or attaches the right answer to the wrong environment. Diagnose each layer, then re-run every pair after a fix. Dependency invalidation removes stale derived answers before release.
Hold or release with evidence
Mark negation and production-block pairs as mandatory. A release passes only if every mandatory prediction matches its reviewed label and no required case is missing. Report operator-level pair consistency, base and edited accuracy, reviewer disputes and source-revision coverage. Keep a visible unresolved state for ambiguous cases rather than forcing an approval. Pairwise reporting explains why the gate failed.
Implementation
def runbook_release_gate(required_cases, expected_decisions, predictions):
missing = sorted(case_id for case_id in required_cases
if case_id not in expected_decisions or case_id not in predictions)
wrong = sorted(case_id for case_id in required_cases
if case_id in expected_decisions and case_id in predictions
and predictions[case_id] != expected_decisions[case_id])
return {"state": "hold" if missing or wrong else "ready",
"missing": missing, "wrong": wrong}
required = {"negation-47", "threshold-82", "paraphrase-91"}
expected = {"negation-47": "blocked", "threshold-82": "blocked",
"paraphrase-91": "approved"}
predicted = {"negation-47": "blocked", "threshold-82": "approved",
"paraphrase-91": "approved"}
assert runbook_release_gate(required, expected, predicted) == {
"state": "hold", "missing": [], "wrong": ["threshold-82"]}
predicted["threshold-82"] = "blocked"
assert runbook_release_gate(required, expected, predicted)["state"] == "ready"
Performance and operating cost
Checking r required cases takes expected O(r) time plus O(r log r) in the worst case for sorted failure IDs, with O(r) space for those lists. Retrieval, answer generation and human adjudication have separate costs. A passing gate covers only the reviewed cases and active source revision; monitor real traffic and repeat the audit after any policy or index change.
Common Mistakes
- Testing only the original passage and omitting its changed counterpart.
- Letting a correct paraphrase score offset a failed mandatory negation case.
- Running contrast cases against an old retrieval index.
- Marking an absent prediction as correct or silently dropping the case.
Read next
- Contrastive text cases: minimal edits and gold-label contracts
- Contrast-pair consistency: sensitivity, invariance and failure types
- Semantic document diffs: changed facts versus wording edits
- Project: release a runbook revision without stale answers
- Generated answers: claim evidence and unanswered state
