This project turns a support team's vague request to 'sort tickets better' into a measurable prompt change. The fictional queue has access, billing, and service-incident labels, plus review when the available text cannot support a choice. Build a zero-shot baseline before adding examples. The deliverable is a task contract, test packet, comparison, and a handoff that names unresolved cases.
Project: establish a support-triage prompt baseline
Write the baseline
Define each label in one sentence and require ticket ID, label, evidence phrase, and uncertainty. Keep the ticket text inside a data section under a fixed instruction envelope. Run 47 labeled development cases, including outage notes that mention invoices and billing disputes that mention service. Record wrong labels by slice, review count, median input tokens, and output completeness. Do not change the label rules while scoring the baseline. A role cue may improve the wording for support agents, but it cannot grant access to refunds or account changes.
Repair one failure class
If the baseline confuses invoice references with billing disputes, add a contrastive pair that isolates the actual customer problem. Hold back 16 independent tickets, including one with an instruction-like subject line and one missing the account identifier needed for a later action. Test the revised prompt against both sets. The classifier may choose a label for the ambiguous account ticket if enough text supports it, but it must not trigger an external action or invent an account ID. Check that every response remains complete under the configured output limit.
Labels: access | billing | incident | review.
Development: 47 labeled tickets; holdout: 16 new tickets.
Baseline: zero-shot task, fixed envelope, required output fields.
Change: one contrastive pair for invoice-versus-outage boundary.
Gate: no wrong-account action; report slice errors and withheld share.
Deliverable: prompt version, case set, results, and unresolved tickets.Performance and operating cost
The baseline uses one call per ticket. Every demonstration adds input tokens to later calls, and the comparison adds another run over the evaluation set. That cost is justified only if the held-out boundary cases improve without regressing ordinary tickets. Validation over the response is linear in output length. Keep the budget for human labeling and adjudication visible; a larger number of cheap model calls does not replace correct expected labels.
Common Mistakes
- Do not add examples until the baseline failure is recorded.
- Do not treat a role cue as permission for a refund.
- Do not remove required fields to meet a short output limit.
Connected lessons
- Zero-shot baseline: measure the task before adding demonstrations
- Prompt envelopes: separate task, evidence, input, and output contract
- Role prompts: use expertise cues without granting authority
- Clarification gates: ask only when a missing fact changes the outcome
- Output budgets: bound length without cutting required facts
- Few-shot prompting: select examples that cover decisions, not just easy cases
- Evaluation sets: measure the failure cases that matter
- Few-shot examples: teach the boundary with near misses
- Prompt foundations decisions
