Resolve an incident event to a service, join that service to the active runbook rule and release an answer only when each hop has reviewed evidence.
Project: audit a multi-step incident answer from event to rule
State the question
An operator asks which rollback rule applies to the service whose retry failed after a second attempt. One log identifies the event; a separate service registry resolves the service; the current runbook supplies the rule. Include a gateway-east distractor, a staging-only rule and an old gateway-west revision. The system must preserve each join rather than selecting a plausible rollback number from memory. Question decomposition names the required facts before retrieval.
Create the evidence fixture
Reviewers mark source spans for the retry event, canonical service ID, environment and active rule. Record document IDs, revisions, access levels and every edge claim. Include a case where the service alias is ambiguous and a case where the current runbook is unavailable; both should stay unanswered. Freeze a set of complete and broken chains before adjusting retrieval or answer prompts. A final-answer-only benchmark cannot distinguish a correct chain from a lucky guess.
Trace and gate the path
Retrieve each step under the policy for its source. Verify the event-to-service link, then the service-to-rule link under the active revision. Reject a chain if any required edge lacks a reviewed source span or if scopes disagree. Evidence-chain checks cover identity and scope; temporal review handles conflicting current and historical statements.
Release a bounded answer
Report per-hop recall, correct joins, complete-chain support, answer correctness and abstention. Inspect every failed path and preserve its evidence IDs for repair. The final response should name its source passage internally and reflect the current rule, while an unresolved case should say which hop could not be verified. The code below enforces the last step on an already reviewed chain; retrieval and reading remain separate tests.
Implementation
def release_chain_answer(required_steps, verified_steps,
answer_value, answer_passage_id):
missing = sorted(set(required_steps) - set(verified_steps))
if missing or not answer_value or not answer_passage_id:
return {"state": "unanswered", "missing_steps": missing}
return {"state": "supported", "answer_value": answer_value,
"answer_passage_id": answer_passage_id}
required = {"retry_event", "service_identity", "active_rule"}
verified = {"retry_event", "service_identity"}
assert release_chain_answer(required, verified, "82 minutes",
"runbook-82") == {
"state": "unanswered", "missing_steps": ["active_rule"]}
verified.add("active_rule")
assert release_chain_answer(required, verified, "82 minutes",
"runbook-82")["state"] == "supported"
Performance and operating cost
Comparing r required steps with verified steps takes expected O(r + v) time and O(r + v) space for set copies, plus O(r log r) sorting when steps are missing. Retrieval and human evidence review dominate end-to-end cost. An apparently complete step list is insufficient unless each step is tied to a valid, current, authorized span.
Common Mistakes
- Returning a memorized rule without resolving the service.
- Calling a chain complete after finding only the event and service.
- Combining evidence from incompatible environments.
- Hiding an unresolved join behind a fluent answer.
Read next
- Multi-step questions: decompose answer slots and dependencies
- Multi-step evidence chains: join identity, revision and access
- Event order: evidence, partial timelines and contradictions
- Project: answer runbook questions with source spans and abstention
- Review claim conflicts across time and source revisions
