Prompt telemetry connects a request to the versioned decision path and its checks while limiting exposure of user data. Record a trace ID, prompt bundle ID, case class, model-call status, validator outcomes, tool-call names, latency, token counts, and effect receipts. Retain raw text only under a defined access and retention policy when it is necessary for investigation. Hashing or redacting a value does not automatically make it harmless; identifiers can still link to people or accounts. Design the event schema before turning on verbose model logging.
Prompt telemetry: measure failures without copying private payloads
Decision in practice
A support assistant processes ticket TK-714 with a customer's payment details in the body. The operational dashboard needs to know that prompt bundle PEB-118 produced a schema failure and one bounded retry. It does not need the card digits or the full ticket text. The event stores the ticket's internal reference, bundle ID, validator failure code, retry count, and latency. A separate protected review store holds a short, redacted sample for an authorized investigator. A test injects a secret-looking value into a tool result and checks that neither ordinary logs nor error traces copy it.
Trace: TR-714; bundle: PEB-118; case_class: billing.
Model calls: 2; validator: schema_missing_field; final_state: review.
Latency: 1840 ms; input/output tokens: 936/127.
Tool effect: none; receipt: none.
Ordinary log excludes ticket body, payment details, and raw tool payload.
Protected sample: redacted, access-controlled, time-limited.Performance and operating cost
An event with fixed fields has O(1) storage per request apart from identifiers and labels; retaining full prompts and responses grows with payload length and can multiply storage and exposure. Aggregating M events is O(M) in the simple case. Redaction, access control, and retention jobs add operational work, so log only fields that support a named diagnostic or audit need. Sample high-volume success traffic while retaining complete counts of failures and external effects. A trace without bundle and case IDs is hard to reproduce; a trace that copies every private field is unnecessarily risky.
Common Mistakes
- Do not enable raw prompt logging by default in a private workflow.
- Do not assume a hash removes all reidentification risk.
- Do not omit effect receipts and validation failures from the trace.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Prompt traces: connect an answer to its inputs, checks, and effects
- Prompt privacy: send only the fields needed for the task
- Observability: join metrics, logs, and traces
- Secrets and configuration across the delivery path
- Prompt changes in CI: test the merge candidate
- Prompt release artifacts: version the whole decision path
- Project: measure a retrieval-backed answer gate
Continue with: Online prompt experiments: define exposure and stop rules first.
Continue with: Sensitive output gates: check the rendered answer before release.
