Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: establish a support-triage prompt baseline

Last updated: 2 Oct 202617 min read
project
IntermediateBy AITrove Editorial

This project turns a support team's vague request to 'sort tickets better' into a measurable prompt change. The fictional queue has access, billing, and service-incident labels, plus review when the available text cannot support a choice. Build a zero-shot baseline before adding examples. The deliverable is a task contract, test packet, comparison, and a handoff that names unresolved cases.

Write the baseline

Define each label in one sentence and require ticket ID, label, evidence phrase, and uncertainty. Keep the ticket text inside a data section under a fixed instruction envelope. Run 47 labeled development cases, including outage notes that mention invoices and billing disputes that mention service. Record wrong labels by slice, review count, median input tokens, and output completeness. Do not change the label rules while scoring the baseline. A role cue may improve the wording for support agents, but it cannot grant access to refunds or account changes.

Repair one failure class

If the baseline confuses invoice references with billing disputes, add a contrastive pair that isolates the actual customer problem. Hold back 16 independent tickets, including one with an instruction-like subject line and one missing the account identifier needed for a later action. Test the revised prompt against both sets. The classifier may choose a label for the ambiguous account ticket if enough text supports it, but it must not trigger an external action or invent an account ID. Check that every response remains complete under the configured output limit.

Output
Labels: access | billing | incident | review.
Development: 47 labeled tickets; holdout: 16 new tickets.
Baseline: zero-shot task, fixed envelope, required output fields.
Change: one contrastive pair for invoice-versus-outage boundary.
Gate: no wrong-account action; report slice errors and withheld share.
Deliverable: prompt version, case set, results, and unresolved tickets.

Performance and operating cost

The baseline uses one call per ticket. Every demonstration adds input tokens to later calls, and the comparison adds another run over the evaluation set. That cost is justified only if the held-out boundary cases improve without regressing ordinary tickets. Validation over the response is linear in output length. Keep the budget for human labeling and adjudication visible; a larger number of cheap model calls does not replace correct expected labels.

Common Mistakes

  • Do not add examples until the baseline failure is recorded.
  • Do not treat a role cue as permission for a refund.
  • Do not remove required fields to meet a short output limit.

Connected lessons

prompt engineering
foundations
Storage details