Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Prompt failure diagnosis: find the broken contract first

Last updated: 5 Oct 202613 min read
tutorial
AdvancedBy AITrove Editorial

Prompt failure diagnosis is the process of identifying which contract broke before changing wording. A bad final answer may begin with a missing task field, stale retrieval result, wrong tool permission, truncated output, or an invalid test label. Freeze the failing request and its permitted evidence first. Record the prompt bundle, model settings, tool schema, retrieval snapshot, validator result, and any external effect receipt. Reproduce with the same inputs before editing the prompt. If the failure cannot be reproduced, preserve the trace and mark the cause unresolved rather than writing a confident story about model behavior.

Classify the observed failure

A support assistant labels ticket TK-934 as billing when the actual request is a service outage with an invoice impact. Inspect the task label definitions and expected decision before adding examples. If the output omits the evidence ID, inspect schema validation and the output ceiling. If it cites a retired policy, inspect retrieval filters and freshness. If it tries to change an account, inspect the server permission decision and tool arguments; no prompt rewrite can supply missing authorization. One response can contain several failures. Record the earliest broken boundary and the downstream symptoms separately so the repair targets the cause.

Run a controlled repair

Choose one hypothesis and change one component at a time: task wording, example pair, evidence selection, parser, tool scope, or model setting. Keep the original case in a development set and preserve an independent holdout. Re-run nearby cases that should remain unchanged. A targeted example may fix the outage-versus-billing boundary while causing the model to overuse the incident label; a shorter output may fix latency while dropping a required field. Compare both errors and costs by slice. If the expected label itself is disputed, adjudicate it before scoring the prompt. Use a negative control to check that the candidate does not produce a forbidden answer merely because its style improved.

Output
Failure TK-934: incident labeled billing.
Bundle: PEB-124; evidence snapshot: IS-83; parser: V5.
Task check: incident wins when outage is the primary event.
Hypothesis: an invoice mention is overweighted.
Change: one contrastive pair in development set only.
Gate: held-out incident slice improves; billing slice does not regress.
Other checks: schema complete, no account effect, cost within budget.

Performance and review cost

Replaying N cases against V variants takes O(NV) model calls, and detailed human review is concentrated on disagreements and high-impact failures. Inspecting an existing trace is often cheaper than generating fresh variants. If several components changed together, the team may need a small factorial experiment or a rollback to isolate the cause; that cost grows rapidly with the number of independent factors. Preserve a compact failure record so future incidents do not pay the same discovery cost. A faster prompt that hides a critical error is not an improvement.

Common Mistakes

  • Do not rewrite prompt prose before checking evidence, parser, and permissions.
  • Do not tune on the final holdout case that should test the repair.
  • Do not call an unresolved label disagreement a model failure.

Wrong task or incomplete answer

Open the checks that match this symptom and inspect their evidence before changing the prompt.

Wrong evidence or unsupported claim

Open the checks that match this symptom and inspect their evidence before changing the prompt.

Malformed or unsafe output

Open the checks that match this symptom and inspect their evidence before changing the prompt.

Unexpected tool behavior

Open the checks that match this symptom and inspect their evidence before changing the prompt.

Regression after a change

Open the checks that match this symptom and inspect their evidence before changing the prompt.

Release follow-through

prompt engineering
field guide
Storage details