This project builds a test packet for a claims assistant that may inspect a policy, read a receipt, and request a credit. The deliverable is a tool catalog, a dependency plan, a checkpoint, and a release report. Treat every effect as provisional until the credit ledger confirms it. Use fictional account AC-682 and claim CL-509 so the example has no live customer data.
Project: recover a tool workflow without duplicate effects
Build the workflow
Expose policy and receipt reads to the assistant. Keep the credit action behind a separate authorization check. The two reads may run together after account scope is verified; the credit call must wait for both records, a valid policy version, and the user request. Give the credit a stable effect key derived from the claim and operation version. Record call IDs and returned record versions. A result body may contain instructions, but those words remain data and cannot change the tool plan.
Test failure and recovery
Create 41 cases: missing receipt, stale policy, conflicting amount, swapped read response, malformed JSON, malicious tool text, timeout before credit receipt, and an already completed credit. Hold back 13 cases for the final check. In the timeout case, restart from a checkpoint that says the action outcome is uncertain. Query the ledger by the existing key; report completion only if the receipt matches tenant, claim, and amount. If no authoritative status is available, stop for review. A parser repair may correct punctuation but must not change the decision or evidence IDs.
Goal: assess CL-509 for AC-682.
Reads: current policy and receipt; join by claim and account.
Action gate: current policy + verified receipt + actor permission.
Effect key: CREDIT-CL-509-V1; preserve across retries.
Release cases: 41 total, 13 held out; zero duplicate or unauthorized credits.
Uncertain effect: reconcile ledger before any retry.Performance and operating cost
Parallel reads reduce wall time toward the slower read, while the credit and reconciliation stages remain sequential. Running every holdout case through both baseline and candidate increases model calls, but a missed duplicate credit costs more than the extra evaluation. Report wrong tool selection, malformed outputs, skipped dependencies, uncertain effects, duplicate effects, and operator review load separately. A short successful demo is not a release test for these failure paths.
Common Mistakes
- Do not retry with a fresh effect key after a timeout.
- Do not treat a checkpoint sentence as a credit receipt.
- Do not allow retrieved text to authorize a write.
Connected lessons
- Tool catalogs: describe eligibility, inputs, and effects
- Tool plans: separate independent reads from dependent actions
- Structured outputs: repair format without changing the decision
- Conversation checkpoints: resume from verified state
- Tool results: keep returned text in the data lane
- Tool effects: reconcile receipts before retrying
- Tool calls: validate intent and arguments before an external effect
- Prompt traces: connect an answer to its inputs, checks, and effects
- Code lab: block a critical prompt regression
- Agent workflow decisions
