Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Annotation prompts: protect holdouts and sample drift

Last updated: 4 Oct 202611 min read
tutorial
AdvancedBy AITrove Editorial

A labeled item should have provenance, rubric version, split assignment, and permitted use. If an evaluation excerpt or its near duplicate appears in a few-shot prompt, the held-out score no longer measures transfer to unseen cases. Freeze split assignment before tuning. In production, sample by time, channel, and uncommon outcomes so a new support policy or product phrase does not disappear in an overall agreement rate. Route changed patterns to human review before relabeling an old test set.

Operational case

Meridian reserves a holdout from the 45 finally labeled excerpts and keeps the two unknowns in a separate review queue. A near-duplicate of one holdout appears in a proposed few-shot prompt, so that example is replaced before evaluation. A later support channel starts using 'swap it' where older tickets said 'send another'. The prompt flags the phrase for rubric review and samples that channel explicitly; it does not retroactively rewrite frozen labels to improve a score.

Output
item: M-28 | rubric: R-47 | split: HOLDOUT
few-shot near-duplicate: reject
unknown queue: 2 items, not negative labels
drift slice: new-channel 'swap it' requests
action: owner review before rubric update

Performance and review cost

Checking N items against H held-out fingerprints is O(N+H) expected work with indexed hashes, while semantic near-duplicate review has a higher human cost. Sampling E live items is O(E) scan work before quotas. Preserve lineage so a prompt revision cannot quietly use its own test answers.

Common Mistakes

  • Do not place held-out excerpts in examples.
  • Do not treat the two unknowns as negative cases.
  • Do not overwrite historical labels to match new policy without versioning.

Connected lessons

prompt engineering
annotation quality
Storage details