An output gate evaluates the generated response against the recipient's verified scope and a field-level disclosure policy before it reaches a page, email, export, or tool. A prompt that says 'do not reveal secrets' is not a sufficient gate. The application must avoid sending unnecessary private data to the model, and it must validate the result independently. Redaction can catch known identifiers and patterns but may miss a paraphrase or expose meaning through context. Keep an allowlist of permitted fields for sensitive workflows, and test the final rendered form, including links, filenames, and metadata.
Sensitive output gates: check the rendered answer before release
Decision in practice
A billing assistant drafts a response for account AC-682. The model includes a correct invoice date but also copies a full bank-routing value from an internal note into the explanation. A simple digit pattern catches that value in the answer body. A second variant hides part of the value in a suggested export filename, which the body-only check misses. The release gate checks the structured response, links, and filename under the account's allowed fields before display. It strips the private field and blocks the export. A reviewer inspects the case without the raw value appearing in routine logs.
Recipient: verified owner of AC-682.
Allowed: invoice date and masked payment method label.
Forbidden: full routing value, other-account data, private note.
Sinks: answer body, link target, filename, export payload.
Gate: reject any forbidden field; log field code, not field value.Performance and operating cost
Scanning L output characters with fixed patterns is O(L), but semantic disclosure checks may require a model or human and raise latency. Build the cheapest reliable gate for known structured fields, then reserve heavier inspection for unstructured high-risk output. The allowlist reduces the set of facts that must be reviewed. A redaction pass is not a replacement for input minimization: once private data enters the model context, it can reappear in unexpected wording. Measure leaks by output sink, not only by chat text.
Common Mistakes
- Do not rely on an instruction alone to prevent disclosure.
- Do not inspect only the visible answer body when links and exports are also outputs.
- Do not log the secret while recording that a redaction gate fired.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Prompt privacy: send only the fields needed for the task
- Generated output: validate again at the destination boundary
- Prompt telemetry: measure failures without copying private payloads
- Request risk routing: classify the action before choosing a response
- Refusal contracts: decline the unsafe effect and preserve useful help
- System prompts: treat instructions as guidance, not a secret vault
- High-impact advice: route uncertain decisions to qualified review
- Project: build a support safety boundary
- Prompt safety decisions
Continue with: API examples: validate payloads and strip credentials.
Continue with: Infrastructure prompts: quarantine state secrets and drift.
Continue with: Email prompts: review To, Cc, Bcc, and reply scope.
