Use auditable rules to propose intake labels, compare them with independent review and keep uncertain tickets in the agent queue.
Project: stage weakly labeled support-ticket triage
Set the decision boundary
At ticket arrival, the service may propose a routing intent; it cannot change a payment, account or customer record. Use only fields available at that moment. Each rule emits a label or abstention, reason and version. Preserve the original message and a record of all votes, including conflicts. The final label may be reviewed and corrected by an agent. Rule functions provide the proposal contract.
Collect and split cases
Choose tickets from several channels and language slices, including negated requests and copied response templates. Group all messages from one customer event before splitting. Have agents independently label a held-out set under one policy. Do not use a queue tag assigned later as either a weak-label source or input feature at intake. The audit boundary screens both rule votes and train/test groups.
Compare three paths
Measure a small manual-only baseline, a rule-vote proposal and a model trained with admitted weak labels. Keep thresholds fixed on a separate validation set. Report per-intent precision, false route rate, abstention, agent queue age and the share of training rows from each rule family. A larger training set is only valuable if independent held-out decisions improve without hiding risk in the reject queue.
Release with traceable rollback
Shadow-route first. Store model, taxonomy, rule set and feature cutoff as one version. Inspect confident wrong proposals as well as conflicts. If a rule begins firing on a new template or a channel’s false route rate rises, disable that rule and return its cases to review while rebuilding affected training rows. A ticket deletion must remove the raw message and any derived training copy.
Implementation
def triage_proposal(rule_votes, allowed_labels):
active = [vote for vote in rule_votes if vote is not None]
if not active:
return {"state": "review", "reason": "no-coverage"}
if any(vote not in allowed_labels for vote in active):
return {"state": "review", "reason": "retired-label"}
if len(set(active)) != 1:
return {"state": "review", "reason": "rule-conflict"}
return {"state": "proposed", "intent": active[0]}
proposal = triage_proposal(["delivery", None, "delivery"],
{"delivery", "refund"})
assert proposal == {"state": "proposed", "intent": "delivery"}
Performance and operating cost
Scanning r votes takes O(r) time and O(r) temporary space. Training on weak labels may reduce manual labeling, but rule maintenance and auditing still cost people-hours. Track cost per correctly routed ticket after review, not the number of generated labels. An abstention that reaches an agent safely is preferable to a wrong automatic route.
Common Mistakes
- Sending a weak label directly to an irreversible action.
- Training on metadata that only exists after agent resolution.
- Claiming success from the volume of newly labeled rows.
- Leaving retired-rule rows in the next training build.
