Build a read-only incident review for a fictional parcel platform. Snapshot PK-214 contains metric MET-54, release diff CFG-8, and redacted support tag set SUP-41. The platform emitted 54 server errors between 09:17 and 09:26 UTC. Release RC-47 at 09:12 changed a connection-pool setting from 32 to 8. The tag set lists 41 contacts in the later window. Those records do not show whether any customer was charged twice. The goal is a grounded incident brief, not a rollback or a proof of causation.
Project: coordinate a parcel-platform incident review
Compare one controller with specialists
Run the same incident packet through a single-agent baseline. Then dispatch a telemetry specialist, release-diff specialist, and redacted-support specialist on named snapshot PK-214. Each packet has its own question, read-only tools, evidence-ID output schema, 45-second deadline, and prohibition on mutations enforced by the runtime. The controller has a two-minute overall deadline. Record model tokens, elapsed time, complete claims, prohibited tool attempts, and unresolved gaps for both paths. Keep the multi-agent workflow only if its case-level gains justify the extra work.
Merge with evidence and version gates
Telemetry can report 54 errors with MET-54. Release review can report the pool change with CFG-8. Support review can report 41 contacts only if it returns the redacted tag set from PK-214. A result from PK-211 is held for a fresh read, even if its prose sounds consistent. Use the claim checker to reject stale packets and unknown evidence identifiers, then inspect whether each cited artifact truly supports its statement. The final brief labels the pool change a candidate mechanism and lists targeted verification, such as pool saturation metrics and a controlled comparison. It states that charge impact is unknown. No specialist can authorize a rollback.
PK-214; RC-47 at 09:12 UTC; 54 errors at 09:17..09:26 UTC.
Telemetry: MET-54; release: CFG-8; support: SUP-41.
Observed: pool setting 32 -> 8; 41 support contacts.
Hypothesis: smaller pool contributed to errors; not yet proven.
Unknown: duplicate charges; no payment evidence.
No writes authorized; late or stale packet stays outside final claims.Performance and operating cost
Three independent specialists can finish near the slowest specialist duration plus merge time, but their token consumption adds. Repeating the same context three times can cost more than one focused run. A timeout triggers a named evidence gap; it does not authorize invented customer impact. Review both final briefs against one held-out rubric covering observed facts, hypothesis language, snapshot alignment, tool scope, and partial-result handling. Retain traces and the final claim ledger so a later correction can be tied to its source packet.
Common Mistakes
- Do not present a before-and-after timeline as causal proof.
- Do not merge a support result from an older snapshot.
- Do not let a specialist with read-only scope deploy a proposed fix.
Connected lessons
- Specialist agents: prove that delegation earns its cost
- Specialist task packets: scope evidence and authority explicitly
- Parallel agents: isolate state and assign one writer
- Specialist synthesis: merge claims, not fluent summaries
- Specialist workflows: bound retries and review the full trace
- Incident triage prompts: build a timestamped evidence ledger
- Prompt release review: assemble the decision packet
- Code lab: reject unknown evidence IDs
- Specialist-agent coordination decisions
