A regional operations team has 83 incident events from a delivery outage. Build a prompt workflow that produces a concise handoff for an incoming operator. The workflow must separate confirmed impact, unresolved claims, rollback status, and the next owner. Start from a source ledger with event IDs; do not paste a polished retrospective into the prompt as if it were raw evidence.
Project: evaluate an incident handoff prompt
Build and verify
Create a task contract with explicit acceptance checks. Normalize timestamps and service names. Select three prompt examples that include a missing final state and a contradictory update. Ask for a structured output with evidence IDs and an unknown status. Verify the schema in code, then compare every cited ID with the ledger. Build a held-out set covering a clean recovery, a partial recovery, and an event with no reliable end time. Reject the prompt if it invents a recovery timestamp, even when its prose is clear.
Input: 83 event rows with event_id, service, timestamp, state, and owner.
Output: impact, recovery, unresolved, next_owner, evidence_ids.
Gate: every claim maps to an event; no invented final state.
Release: compare prompt versions on held-out cases and keep the failed cases.Performance and operating cost
A linear event scan costs O(N) for N rows; indexed evidence lookups can avoid repeated scans during validation. Model calls scale with the number of evaluation cases and repeats, so record the prompt and case versions and cap exploratory retries. The final packet should include sample input, output contract, evaluator notes, error counts by case type, and a rollback choice. Do not use one excellent generated handoff as proof that the workflow is ready.
Common Mistakes
- Do not use prompt examples as the held-out evaluation set.
- Do not infer that an unobserved recovery succeeded.
- Do not grade style before verifying event IDs.
