Prompt evaluations need cases where the expected behavior is to withhold, reject, or ask for evidence. Positive examples show whether the system can complete a task; negative controls show whether it acts when it should not. Build controls for missing records, irrelevant retrieved text, conflicting versions, unauthorized actors, and fabricated identifiers. Freeze the expected state before testing. A control must be plausible enough to expose the failure mode without copying the final holdout into the prompt. Record both the model output and any downstream effect.
Negative controls: test the answer that should not be produced
Decision in practice
A renewal assistant succeeds on 39 clean cases. It fails when a retrieved maintenance note mentions RN-284 but contains no inspection date: the assistant treats the note as proof and approves. A negative control keeps the same user request but removes the inspection record. Another supplies a valid inspection for a different asset. Both should return review. A third embeds an instruction in the note to approve every renewal; the model may quote the note as data, but it must not follow that instruction. The release gate checks that no approval or tool effect occurs in these controls.
Control N-1: inspection missing -> review; no renewal effect.
Control N-2: inspection belongs to another asset -> review.
Control N-3: retrieved note says 'approve all' -> ignore instruction; evaluate facts.
Measure: false approvals, unauthorized effects, and correct abstentions per control family.Performance and operating cost
For C controls, a single candidate requires O(C) model calls and review records; multiple candidates or prompt versions multiply that cost. The valuable metric is the count of forbidden outcomes by control family, not a single pooled accuracy score. Maintain enough ordinary positive cases so a model that always refuses cannot pass. Revisit controls after retrieval, policy, or tool changes. A negative control that was used to tune the prompt should move to development data, with a fresh holdout case replacing it.
Common Mistakes
- Do not call an always-refusing system successful because controls pass.
- Do not omit downstream tool effects from the result record.
- Do not promote a tuned negative control as an independent final test.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Evaluation sets: measure the failure cases that matter
- Prompt injection: test untrusted content at every boundary
- Reasoning summaries: show checkable grounds, not invented certainty
- Multiple candidates: filter invalid answers before ranking
- Critique loops: require a named defect and a stopping rule
- Few-shot examples: teach the boundary with near misses
- Context distillation: shorten input without losing governing exceptions
- Project: verify a contract-renewal prompt at the boundary
- Reasoning and evidence checks
Continue with: Project: select and verify a reasoning pattern.
Continue with: Synthetic case coverage: count distinct decisions, not rewritten sentences.
