Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Evaluation leakage: keep the release test independent

Last updated: 5 Oct 202610 min read
tutorial
AdvancedBy AITrove Editorial

A release evaluation needs cases independent of the material used to write the prompt, choose examples, and tune settings. Deduplicate near-identical records across development and holdout sets, including paraphrases and versions of the same incident. Freeze expected labels before running the candidate. If a failure becomes a prompt example, move that specific case out of the final holdout and add a fresh case of the same failure class. Preserve a hidden or later-collected set for the final check.

Decision in practice

A support team has 92 labeled tickets. An author copies six troublesome holdout tickets into the few-shot prompt, then reports a large gain on the same 92. The gain cannot establish generalization. The evaluator marks those six and their paraphrases as contaminated, creates a new holdout from later tickets, and compares the revised prompt with the previous version there. It also verifies that the expected answers were not embedded in a retrieved knowledge page. The team retains the contaminated cases for debugging, but stops calling them independent release evidence.

Output
Development set: prompt editing and example selection.
Frozen holdout: unseen cases and labels.
Leak check: duplicate IDs, near-duplicate content, retrieved answer keys.
If contaminated: replace release cases; keep old ones for diagnosis only.

Performance and operating cost

Deduplication takes additional data work and may require semantic review when cases are paraphrased. A simple hash finds exact duplicates in O(N) expected time with an index, but near-duplicates need comparison or retrieval. The cost matters because a misleading evaluation can ship a broken workflow. Log the lineage of every test case and sample example. Repeat the final check on fresh data after a policy change rather than recycling a set that authors have memorized.

Common Mistakes

  • Do not grade on the examples inside the prompt.
  • Do not treat paraphrased duplicates as independent cases.
  • Do not hide contaminated cases; label and retain them for debugging.

Connected lessons

Continue with: Evaluation case ledgers: revise labels without erasing history.

Continue with: Synthetic evaluation holdouts: stop the generator from teaching the test.

Continue with: Forecast prompts: enforce the as-of data boundary.

Continue with: Annotation prompts: protect holdouts and sample drift.

Apply this boundary: Group and time validation: split by the failure you expect in production.

prompt engineering
tutorial
Storage details