Rules derived from text patterns or metadata can accelerate labeling, but they produce noisy proposals rather than reviewed truth.
Weak label functions: coverage, conflict and abstention
Define each rule as a source
A label function might recognize a refund phrase or a routing queue tag. It should return one taxonomy label or abstain, plus a reason and rule version. Do not force a default label when its evidence is absent. Separate functions built from customer text from functions built from downstream agent actions; the latter may encode the answer and leak future information into training. The label policy defines the allowed target labels.
Measure the label matrix
Compute each function’s coverage, overlap and conflicts on the unlabeled pool, then compare its precision on a small independently reviewed set. A rule that fires on nearly everything may inflate coverage while adding no useful signal. Correlated rules are not independent votes: two phrases copied from one template may make the same mistake. Retain the full function outputs with source IDs so an error can be traced to a rule, not just a final consensus label.
Keep abstention useful
Some records will receive no vote or incompatible votes. Route a sample to human review, especially high-cost intents. An automatic tie-breaker based on alphabetical label order is not defensible. Track unresolved coverage by language, channel and time, since a stable overall rate can hide a new cohort with no useful functions. Conflict audits decide whether a training label is admissible.
Separate training labels from evaluation
Never evaluate against labels generated by the same functions used to train. Freeze a human-reviewed audit and group related conversations before splitting. Compare a manual-label baseline, weak labels and a mixed system on that audit. Report false automation and review burden, not merely how many extra rows were labeled. The ticket project applies these gates.
Implementation
def combine_rule_votes(votes):
active = [label for label in votes if label is not None]
if not active:
return None, "no-coverage"
unique = set(active)
if len(unique) > 1:
return None, "conflict"
return active[0], "unanimous-proposal"
assert combine_rule_votes(["refund", None, "refund"]) == (
"refund", "unanimous-proposal")
assert combine_rule_votes(["refund", "delivery"])[1] == "conflict"
Performance and operating cost
Combining r votes is O(r) time and O(r) space in this straightforward implementation. More elaborate label models add fitting cost and assumptions about rule dependence. The main saving is reviewer time, but that saving disappears if noisy training labels create wrong production routes. Sample both unanimous and conflicting cases for audit.
Common Mistakes
- Treating rule output as human-reviewed ground truth.
- Counting correlated template rules as independent evidence.
- Assigning a default label to every abstention.
- Evaluating on labels created by the same rules.
