A routing metric needs class counts, language and channel slices, and a policy for cases the model should send to human review.
Text classification evaluation: inspect slices and allow abstention
Tie the metric to work
A false urgent route consumes specialist capacity; a missed urgent ticket delays response. Set a primary metric tied to those costs, then show a confusion matrix and per-class support. An overall score can rise while the rare urgent queue becomes worse. Threshold choice] should use a validation set and a measured review budget.
Inspect relevant slices
Split results by language, product, message length and arrival channel. State numerator and denominator for every slice, and suppress or qualify unstable estimates from tiny counts. A model can perform well on long English descriptions while failing on short mixed-language messages. The slice list should be written before reading the final test set.
Make uncertainty actionable
A route confidence below a declared threshold can go to manual triage. An abstention is not automatically a success; count its volume, delay and eventual correct destination. If the classifier cannot represent a new queue, an out-of-scope detector or fallback is needed. Serving] must expose the fallback consistently.
Test operational capacity
Simulate a week of tickets with a realistic class mix. At each threshold, count automated routes, manual reviews, urgent misses and wrong specialist assignments. Compare the chosen policy with a simple rule and with no automation. Re-evaluate after a queue taxonomy change rather than assuming old labels retain their meaning.
Implementation
def route_with_abstention(class_scores, class_names, minimum_score):
if len(class_scores) != len(class_names) or len(class_scores) == 0:
raise ValueError("invalid score contract")
best_index = max(range(len(class_scores)), key=class_scores.__getitem__)
if class_scores[best_index] < minimum_score:
return "manual-review"
return class_names[best_index]Performance and operating cost
Threshold selection over N scored tickets costs O(N) for each candidate threshold; sorting once supports a sweep in O(N log N). Manual-review volume is an operational cost and must be budgeted.
Common Mistakes
- Do not hide rare-class failures in one average.
- Do not count abstention as a correct automated route.
- Do not choose thresholds on the untouched final test set.
