Parse procedural text into reviewable action frames while keeping prerequisites, warnings and quoted evidence separate from executable steps.
Runbook steps: actions, targets and source conditions
A verb phrase is not an instruction to execute
“Restart gateway-west after confirming replica health” contains an action, target and condition. “Do not restart gateway-west” contains a prohibition. Both mention the same action, but they have opposite operational meanings. Extract source spans and classify each clause as step, prerequisite, warning, observation or prohibition. Do not turn a detected verb into a tool call. Predicate roles and negation help bind the right target and polarity.
Keep conditions attached
Runbooks often place prerequisites in earlier bullets or sections: drain traffic, confirm standby readiness, then restart a service. A single sentence may not contain all conditions. Track explicit links between steps and preserve any unresolved reference such as “once this completes.” Separate observed postconditions from required preconditions. A health check can be an instruction or an observed result depending on tense and section. Document sections help keep headings and warnings in scope.
Use a constrained action vocabulary
Define approved action types and target kinds in an operator-owned schema. A model may suggest a plausible operation absent from the runbook; that suggestion should be rejected. Keep original text beside the normalized action frame and require the span to match the source revision. The example below checks an extracted frame’s shape, source text and allowed action. It does not decide whether the action is safe or authorized.
Evaluate the dangerous confusions
Measure step detection, action-target links, condition recall, warning preservation and false executable steps. Include negated instructions, quoted examples, obsolete sections and steps that refer to an earlier result. An extractor that finds every restart mention but loses “do not” is unsafe. The dependency lesson checks ordering, and the project tests an operator review flow.
Implementation
def validate_step_frame(source_text, frame, approved_actions):
start, end = frame["source_span"]
if not 0 <= start < end <= len(source_text):
return {"state": "review", "reason": "invalid-span"}
if source_text[start:end] != frame["source_phrase"]:
return {"state": "review", "reason": "span-mismatch"}
if frame["kind"] != "step" or frame["action"] not in approved_actions:
return {"state": "review", "reason": "not-approved-step"}
if not frame.get("target_id") or not frame.get("condition_span"):
return {"state": "review", "reason": "missing-binding"}
return {"state": "frame-valid"}
source = "After replica health is green, restart gateway-west."
start = source.index("restart")
frame = {"source_span": (start, len(source)),
"source_phrase": source[start:], "kind": "step",
"action": "restart", "target_id": "gateway-west",
"condition_span": (0, start)}
assert validate_step_frame(source, frame, {"restart"})["state"] == "frame-valid"
assert validate_step_frame(source, {**frame, "kind": "warning"},
{"restart"})["state"] == "review"
Performance and operating cost
The fixed-field checks are O(e) time to compare e evidence characters and O(e) for the source slice in Python. Full parsing scales with runbook length and model cost. A valid frame only means the proposed action is grounded and shaped correctly; policy and operator authorization remain separate.
Common Mistakes
- Treating every imperative verb as executable.
- Dropping “do not” from an extracted step.
- Losing a prerequisite in the preceding section.
- Accepting a model-invented action not present in the approved vocabulary.
