Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Classification prompts: write the label boundary first

Last updated: 5 Oct 202611 min read
tutorial
AdvancedBy AITrove Editorial

A classification prompt should provide a finite label set, the evidence required for each label, precedence or multi-label rules, and an explicit unknown route. Labels should represent business decisions that can be reviewed, not vague sentiment about the text. Include near-boundary examples in evaluation, especially cases that mention a topic without requesting its associated action. The model's label is a proposal; downstream effects still require separate authorization. Track disagreements by case type rather than optimizing only overall accuracy, which can hide a rare but costly class.

Operational case

A support queue uses labels delivery_delay, damaged_item, billing_dispute, and unknown. Ticket TK-917 says a package arrived late and the outer box was torn, but the customer asks only when a replacement for the damaged item will ship. If the queue permits multiple labels, the case can carry both delivery_delay and damaged_item while damaged_item owns the response workflow. Ticket TK-918 merely asks what the late-delivery policy says; it is not evidence that this customer's parcel was late. A thin prompt that keys on the word late would misroute the second ticket. The rubric requires a claimed event, requested action, and evidence span.

Output
TK-917: claimed late arrival + torn box; request=replacement status.
Labels: delivery_delay, damaged_item; primary=damaged_item.
TK-918: policy question only; no claimed late event.
Route: informational policy answer, not delay compensation.
Unknown: missing event or conflicting account evidence -> human review.

Performance and operating cost

A fixed rubric over L labels has O(L) decision work per case in a simple implementation; review time rises with ambiguous labels and overlapping policies. Adding more examples increases prompt tokens and can overfit to wording rather than the intended boundary. Use a held-out set with positive, negative, and near-boundary cases for each label. Report per-class precision and recall and the share sent to unknown. A classifier that rarely abstains may appear efficient while passing unsupported cases into an automated effect.

Common Mistakes

  • Do not classify a policy question as a reported incident.
  • Do not force an uncertain ticket into a known class to avoid an unknown label.
  • Do not use the label alone as approval for a refund or replacement.

Connected lessons

Continue with: Meeting minutes: separate decisions from proposals.

Continue with: Support prompts: route by issue and verified urgency.

Continue with: Annotation prompts: define the unit and label taxonomy.

prompt engineering
data workflows
Storage details