Ship a feedback pipeline that separates delivery, billing and support opinions, keeps uncertain tuples for review and reports what the model misses.
Project: turn mixed product feedback into reviewed aspect signals
Write the product question
A weekly product review needs to know which feature caused complaints, not whether a whole ticket sounds unhappy. Define three initial aspects—delivery, billing and support—and permit “other” when none fits. An output is a set of target-polarity tuples with source offsets, model version and review state. It is not a customer-risk score. Establish who may inspect raw text and how long examples are retained.
Build a reviewed corpus
Capture original customer messages, exclude agent templates from customer opinion labels and group all replies from one case. Annotate target spans, opinion cues and negation scope; adjudicate mixed statements and quotations. Freeze a later-time audit set with at least one sufficiently reviewed sample for each aspect and language mix. Where support is thin, state that the slice is unverified. The target and negation policy supplies the annotation contract.
Compare and gate systems
Start with a sparse text baseline for each aspect, then test a contextual candidate on the same split. Measure exact tuple F1, severe-complaint recall, incorrect target rate and manual-review volume. Calibrate scores on validation cases and freeze thresholds before final audit. Route unreviewed aspects to review. The calibration lesson turns scores into safe decisions; multilingual routing covers related fallback design.
Operate the weekly report
Publish counts only for reviewed or accepted tuples and show sample size beside every percentage. A changed product taxonomy requires relabeling, not merely renaming a chart. Monitor corrections from a random sample of high-confidence results and from abstentions. Keep raw examples in a restricted review queue, not in the general dashboard. Compare week-to-week movement only after checking whether traffic mix, channel or label policy changed.
Implementation
def weekly_aspect_counts(decisions):
counts = {}
for decision in decisions:
if decision["state"] not in {"reviewed", "automatic"}:
continue
key = (decision["aspect"], decision["polarity"])
counts[key] = counts.get(key, 0) + 1
return counts
feedback = [
{"aspect": "delivery", "polarity": "negative", "state": "reviewed"},
{"aspect": "support", "polarity": "positive", "state": "automatic"},
{"aspect": "billing", "polarity": "negative", "state": "manual-review"},
]
assert weekly_aspect_counts(feedback) == {
("delivery", "negative"): 1, ("support", "positive"): 1
}
Performance and operating cost
The aggregation is O(n) time and O(a·p) space in the worst case for n decisions, a aspects and p polarities. Model inference and review are the recurring costs; report processing latency and estimated cases awaiting review. Weekly counts are not population estimates if the accepted set is selected by a threshold. Present coverage and traffic mix with the counts to avoid false trend claims.
Common Mistakes
- Publishing a change in complaint percentage without checking traffic composition.
- Counting unreviewed low-confidence tuples as facts.
- Treating customer text and copied agent replies as the same speaker.
- Assuming a product taxonomy change preserves historical comparability.
