Extract a maintenance runbook into sourced step frames and dependency edges, then hold prohibited, cyclic or unverified plans for operator review.
Project: audit maintenance runbook steps before execution
Prepare the corpus
Create fictional runbook revisions for gateway-west: verify replica health, drain traffic, restart the gateway and check recovery. Add a warning saying not to restart before the drain is confirmed. Include a quoted command example that must not become an action, a missing prerequisite and an obsolete revision. Annotate step spans, targets, conditions, warnings and dependency edges. Step frames define the atomic labels.
Build a reviewable graph
Parse document structure, detect candidate steps, bind action and target, attach explicit conditions and order edges, then preserve each source span and revision. Check for missing step IDs and cycles. A model may propose a dependency, but the operator should see the text that supports it. Do not silently generate a command from a paraphrased step. Graph checks only establish an ordering, not a safe execution plan.
Hold at the action boundary
Require an approved action vocabulary, current runbook revision, verified conditions and an authorized operator decision before a separate system may act. The code below returns review-ready only when a prepared record clears structural gates; it does not execute or authorize a restart. A warning or unverified condition keeps the record on hold. Keep the last verified observation and its timestamp next to each condition.
Evaluate failure modes
Report action-target accuracy, condition recall, false executable steps, dependency cycles found, stale-revision detection and false-ready rate. Inspect each mistake involving prohibitions or prerequisites. Compare output with operator-reviewed gold graphs, and run regression checks after runbook edits. If evidence is incomplete, an ordered graph can still be useful for human review while execution remains blocked.
Implementation
def admit_runbook_review(plan, current_revision):
if plan.get("revision") != current_revision:
return {"state": "hold", "reason": "stale-revision"}
if plan.get("cycle") or plan.get("missing_step"):
return {"state": "hold", "reason": "graph-invalid"}
if plan.get("prohibited_action"):
return {"state": "hold", "reason": "prohibition"}
if any(state != "verified" for state in plan.get("conditions", {}).values()):
return {"state": "hold", "reason": "condition-unverified"}
if not plan.get("operator_id") or not plan.get("source_spans"):
return {"state": "hold", "reason": "missing-review-context"}
return {"state": "review-ready", "operator_id": plan["operator_id"]}
plan = {"revision": "runbook-r47", "cycle": False,
"missing_step": False, "prohibited_action": False,
"conditions": {"replica-health": "verified"},
"operator_id": "operator-82", "source_spans": [(12, 47)]}
assert admit_runbook_review(plan, "runbook-r47")["state"] == "review-ready"
assert admit_runbook_review({**plan, "prohibited_action": True},
"runbook-r47")["state"] == "hold"
Performance and operating cost
For c condition states, this final gate takes O(c) time and O(1) extra space. Parsing and graph validation cost more and depend on document size. Review-ready is an editorial state only; execution needs a separate trusted service, fresh observations and explicit authorization.
Common Mistakes
- Running a command because a model extracted an action phrase.
- Ignoring a negated warning in another section.
- Using stale observations to satisfy a prerequisite.
- Treating a cycle-free order as proof the runbook is safe.
