A prompt trace records the request ID, prompt and model versions, input and evidence IDs, retrieval decisions, parser result, tool proposals, executor verdicts, and final status. It need not copy the full private request. Use redacted fields and scoped access, and retain traces according to the product policy. Distinguish a generation failure from a retrieval gap, schema failure, tool denial, and human-review wait. Those failure classes lead to different fixes.
Prompt traces: connect an answer to its inputs, checks, and effects
Decision in practice
A support team sees a jump in incorrect warranty denials. Traces show the prompt version stayed fixed, but the retrieval index began returning a retired coverage document. The team reindexes the policy and replays affected case IDs. It does not rewrite prompt prose to compensate for stale input. For a separate incident, traces show valid evidence but a parser rejected an omitted enum; that case belongs to output-contract repair. Reviewers can audit which cases were displayed and which caused a write.
Trace: request_id, template_hash, model_id, evidence_ids, index_version.
Checks: schema_result, business_rule_result, tool_policy_result.
Outcome: displayed, escalated, or effect_receipt_id.
Privacy: redacted payloads and access-controlled retention.Performance and operating cost
Tracing consumes storage roughly proportional to request volume and recorded metadata. Full prompt capture may be inappropriate for private inputs, so store hashes and source IDs where replay can retrieve data under authorization. Monitor error rates by failure class and version, not just total requests. Sampling can reduce telemetry cost for routine successes, but keep complete records for external effects and critical failures. A trace is operational evidence, not a substitute for evaluating answer correctness.
Common Mistakes
- Do not log secrets to make debugging easier.
- Do not blame prompt wording for a stale retrieval index.
- Do not mix rejected tool proposals with executed effects in one counter.
Connected lessons
- Prompt Engineering
- Production prompt engineering
- Prompt releases: version the whole decision path and keep a rollback
- Evidence IDs: make generated claims auditable against supplied records
- Human handoff: preserve evidence and the reason for uncertainty
- Model migration: replay contracts before changing providers or versions
- Project: regression-test a customer triage prompt
- Advanced prompt engineering decisions
Related implementation
Field guide: Prompt failure diagnosis: find the broken contract first.
Continue with: Incident triage prompts: build a timestamped evidence ledger.
Continue with: Specialist workflows: bound retries and review the full trace.
Continue with: Recurring prompts: make every run replayable and auditable.
