Prompt failure diagnosis is the process of identifying which contract broke before changing wording. A bad final answer may begin with a missing task field, stale retrieval result, wrong tool permission, truncated output, or an invalid test label. Freeze the failing request and its permitted evidence first. Record the prompt bundle, model settings, tool schema, retrieval snapshot, validator result, and any external effect receipt. Reproduce with the same inputs before editing the prompt. If the failure cannot be reproduced, preserve the trace and mark the cause unresolved rather than writing a confident story about model behavior.
Prompt failure diagnosis: find the broken contract first
Classify the observed failure
A support assistant labels ticket TK-934 as billing when the actual request is a service outage with an invoice impact. Inspect the task label definitions and expected decision before adding examples. If the output omits the evidence ID, inspect schema validation and the output ceiling. If it cites a retired policy, inspect retrieval filters and freshness. If it tries to change an account, inspect the server permission decision and tool arguments; no prompt rewrite can supply missing authorization. One response can contain several failures. Record the earliest broken boundary and the downstream symptoms separately so the repair targets the cause.
Run a controlled repair
Choose one hypothesis and change one component at a time: task wording, example pair, evidence selection, parser, tool scope, or model setting. Keep the original case in a development set and preserve an independent holdout. Re-run nearby cases that should remain unchanged. A targeted example may fix the outage-versus-billing boundary while causing the model to overuse the incident label; a shorter output may fix latency while dropping a required field. Compare both errors and costs by slice. If the expected label itself is disputed, adjudicate it before scoring the prompt. Use a negative control to check that the candidate does not produce a forbidden answer merely because its style improved.
Failure TK-934: incident labeled billing.
Bundle: PEB-124; evidence snapshot: IS-83; parser: V5.
Task check: incident wins when outage is the primary event.
Hypothesis: an invoice mention is overweighted.
Change: one contrastive pair in development set only.
Gate: held-out incident slice improves; billing slice does not regress.
Other checks: schema complete, no account effect, cost within budget.Performance and review cost
Replaying N cases against V variants takes O(NV) model calls, and detailed human review is concentrated on disagreements and high-impact failures. Inspecting an existing trace is often cheaper than generating fresh variants. If several components changed together, the team may need a small factorial experiment or a rollback to isolate the cause; that cost grows rapidly with the number of independent factors. Preserve a compact failure record so future incidents do not pay the same discovery cost. A faster prompt that hides a critical error is not an improvement.
Common Mistakes
- Do not rewrite prompt prose before checking evidence, parser, and permissions.
- Do not tune on the final holdout case that should test the repair.
- Do not call an unresolved label disagreement a model failure.
Wrong task or incomplete answer
Open the checks that match this symptom and inspect their evidence before changing the prompt.
- Prompt engineering: define a task that can be checked
- Prompt intent: turn an open request into an acceptance check
- Clarification gates: ask only when a missing fact changes the outcome
- Output budgets: bound length without cutting required facts
Wrong evidence or unsupported claim
Open the checks that match this symptom and inspect their evidence before changing the prompt.
- Retrieved context: select sufficient evidence before writing the answer
- Retrieved evidence: reconcile versions and conflicting facts
- Evidence IDs: make generated claims auditable against supplied records
- Cross-modal evidence: keep conflicting observations separate
Malformed or unsafe output
Open the checks that match this symptom and inspect their evidence before changing the prompt.
- Output contracts: parse a result and preserve an explicit unknown state
- Structured outputs: repair format without changing the decision
- Generated output: validate again at the destination boundary
- Sensitive output gates: check the rendered answer before release
Unexpected tool behavior
Open the checks that match this symptom and inspect their evidence before changing the prompt.
- Tool calls: validate intent and arguments before an external effect
- Tool results: keep returned text in the data lane
- Tool effects: reconcile receipts before retrying
- Request risk routing: classify the action before choosing a response
Regression after a change
Open the checks that match this symptom and inspect their evidence before changing the prompt.
- Paired prompt evaluation: count changes, then inspect uncertainty
- Evaluation labels: adjudicate disagreement before scoring a release
- Prompt release artifacts: version the whole decision path
- Prompt traces: connect an answer to its inputs, checks, and effects
