Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Zero-shot baseline: measure the task before adding demonstrations

Last updated: 5 Oct 20269 min read
tutorial
BeginnerBy AITrove Editorial

A zero-shot prompt states the task, input, constraints, and output shape without showing task-specific input-output pairs. It gives an inexpensive baseline for deciding whether examples are needed. Record the model setting and run a representative case set before changing wording. If the baseline fails on one class of ambiguity, add a small example that isolates that boundary and compare against the same cases. Do not add demonstrations merely because the technique is popular; examples consume context and can import their own errors or private details.

Decision in practice

A support desk sorts requests into access, billing, or service incident. The first prompt defines the three labels, allows review when no label fits, and requires one ticket ID in the output. On 47 held-out tickets it handles clear requests but misclassifies notes that mention an invoice during a service outage. The team records that slice before adding any example. It later adds one paired demonstration that separates an invoice dispute from an outage with billing impact, then reruns the original cases and a fresh holdout. The gain must appear on unseen tickets rather than on the demonstration itself.

Output
Task: label ticket TK-714 as access, billing, incident, or review.
Input: ticket text and ID; no customer payment details.
Output: label, ticket_id, short evidence phrase.
Boundary: an outage that mentions an invoice remains an incident.
Baseline: run 47 frozen cases before adding examples.

Performance and operating cost

One zero-shot call per case gives N calls for N cases, while every added demonstration increases input tokens across later requests. For E examples of T average tokens, the prompt grows by roughly O(ET) tokens per call. Evaluation review remains the main setup cost. Compare error counts by case type, withheld share, and latency, not a single attractive example. If a tiny prompt already passes the acceptance set, extra examples may add expense without benefit.

Common Mistakes

  • Do not treat one successful response as a baseline.
  • Do not add examples before identifying the failing case class.
  • Do not use a final holdout ticket as a demonstration.

Connected lessons

prompt engineering
foundations
Storage details