A release evaluation needs cases independent of the material used to write the prompt, choose examples, and tune settings. Deduplicate near-identical records across development and holdout sets, including paraphrases and versions of the same incident. Freeze expected labels before running the candidate. If a failure becomes a prompt example, move that specific case out of the final holdout and add a fresh case of the same failure class. Preserve a hidden or later-collected set for the final check.
Evaluation leakage: keep the release test independent
Decision in practice
A support team has 92 labeled tickets. An author copies six troublesome holdout tickets into the few-shot prompt, then reports a large gain on the same 92. The gain cannot establish generalization. The evaluator marks those six and their paraphrases as contaminated, creates a new holdout from later tickets, and compares the revised prompt with the previous version there. It also verifies that the expected answers were not embedded in a retrieved knowledge page. The team retains the contaminated cases for debugging, but stops calling them independent release evidence.
Development set: prompt editing and example selection.
Frozen holdout: unseen cases and labels.
Leak check: duplicate IDs, near-duplicate content, retrieved answer keys.
If contaminated: replace release cases; keep old ones for diagnosis only.Performance and operating cost
Deduplication takes additional data work and may require semantic review when cases are paraphrased. A simple hash finds exact duplicates in O(N) expected time with an index, but near-duplicates need comparison or retrieval. The cost matters because a misleading evaluation can ship a broken workflow. Log the lineage of every test case and sample example. Repeat the final check on fresh data after a policy change rather than recycling a set that authors have memorized.
Common Mistakes
- Do not grade on the examples inside the prompt.
- Do not treat paraphrased duplicates as independent cases.
- Do not hide contaminated cases; label and retain them for debugging.
Connected lessons
- Prompt Engineering
- Production prompt engineering
- Few-shot prompting: select examples that cover decisions, not just easy cases
- Evaluation sets: measure the failure cases that matter
- Selective answers: measure when to abstain
- Generated output: validate again at the destination boundary
- Project: measure a retrieval-backed answer gate
- Prompt evidence and output decisions
Continue with: Evaluation case ledgers: revise labels without erasing history.
Continue with: Synthetic evaluation holdouts: stop the generator from teaching the test.
Continue with: Forecast prompts: enforce the as-of data boundary.
Continue with: Annotation prompts: protect holdouts and sample drift.
Apply this boundary: Group and time validation: split by the failure you expect in production.
