A high polarity score can still be wrong for rare aspects, mixed-language feedback or copied text. Measure the tuple decision in each reviewed slice.
Aspect sentiment calibration across product and language slices
Separate model confidence from routing confidence
An aspect extractor may identify “delivery” with high confidence while a polarity model hesitates. A joint action should depend on both, not the larger of the two. Calibrate the probability of a correct aspect-polarity tuple on held-out reviewed cases. A score learned from popular product features may not transfer to a new feature with only a few labels. For unsupported slices, abstain and collect review rather than silently borrowing a threshold from the largest slice.
Choose the split before thresholds
Keep customer conversations and copied reviews together, and reserve later messages for final audit. Fit calibration on validation data only. For each candidate threshold, count accepted tuples, correct accepted tuples, missed negative feedback and review volume. A different threshold for every tiny segment will overfit; group sparse slices under one conservative fallback. Grouped temporal validation protects the audit from duplicate language.
Inspect costly failures
Product teams usually care about a missed severe complaint more than a mild false alarm, but the cost must be written down rather than guessed. Report per-aspect recall, wrong-target assignments, false negative complaints, automatic coverage and calibration error with support counts. Split multi-aspect, negated, sarcastic, translated and Romanized messages. Script profiling can describe input mix; it cannot provide the reviewed language label.
Turn scores into a decision contract
Return tuples with target offsets and a decision state: automatic, reviewed-needed or unsupported. Record the model, annotation, threshold and tokenizer versions together. If an aspect has no reviewed support, route it to a human queue. Sample some high-confidence automatic results for post-release review, or selective labels will hide confident errors. The contract extends classification abstention to multi-target text and feeds the applied feedback workflow.
Implementation
def aspect_decision(aspect_score, polarity_score, aspect_threshold,
polarity_threshold, reviewed_slice):
if reviewed_slice is None:
return "manual-review", "unreviewed-slice"
if not all(0 <= score <= 1 for score in (aspect_score, polarity_score)):
raise ValueError("scores must be probabilities")
if aspect_score < aspect_threshold:
return "manual-review", "uncertain-target"
if polarity_score < polarity_threshold:
return "manual-review", "uncertain-polarity"
return "automatic", "both-gates-passed"
assert aspect_decision(0.93, 0.61, 0.82, 0.74, "billing-reviewed") == (
"manual-review", "uncertain-polarity"
)
Performance and operating cost
The decision gate is O(1) time and space per tuple. Calibration fitting, model inference and manual review dominate total cost; estimate review cases as arrival volume times abstention rate for each slice. High selective accuracy with low coverage may simply move work to humans. Report both figures and the reviewed support count before claiming a release improved operations.
Common Mistakes
- Using the aspect score as a proxy for polarity correctness.
- Tuning thresholds on the final audit set.
- Assigning an unsupported aspect the English-majority threshold.
- Reporting accepted-case accuracy while hiding review volume and missed complaints.
Read next
- Aspect sentiment: bind opinions to targets and negation
- Project: turn mixed product feedback into reviewed aspect signals
- Text classification evaluation: inspect slices and allow abstention
- Text validation: split conversations, duplicates and time together
- Calibrate multilingual text decisions and fallback routes
