The same template can trigger several rules and appear in both train and test. Audit these dependencies before claiming a gain.
Weak-label audits: correlated rules and split leakage
Trace every vote to evidence
A weakly labeled row needs rule IDs, rule revisions, source fields, conflict state and the text revision they saw. A billing keyword and a billing queue tag may both point to “refund,” but the queue tag might have been set after the final outcome. Exclude future metadata from a model that will predict at message arrival. Keep an explicit feature cutoff timestamp. The function contract defines a vote and abstention.
Audit correlation and error
Inspect rule overlap and shared failure cases on reviewed data. If several rules are variants of one template, their agreement does not multiply confidence. Measure precision for each rule, for the combined proposal and for the abstained set. Review high-impact conflicts as cases, preserving disagreement reasons. A rule with strong overall precision may fail on code-switched text or negated requests; report those slices rather than hiding them in a pooled average.
Build a clean holdout
Group messages by customer event, incident and reused template before splitting. Remove synthetic descendants and near duplicates from the opposite split. The final audit must be labeled independently of the weak functions, with a frozen policy version. If the same downstream queue tag generated labels and appears as an input feature, the validation is contaminated even with separate tickets. Grouped temporal validation supplies the broader split design.
Gate retraining and rollback
Version each function and its source schema. A new queue process or message template changes rule behavior, so monitor coverage, conflicts and reviewer precision after release. Retrain only when the new label matrix and independent audit both pass. Keep the previous model and label set so a bad rule update can be reversed. The applied project makes these checks operational.
Implementation
def weak_label_row_allowed(row, arrival_time):
if row["text_revision"] != row["rule_text_revision"]:
return False
if row["feature_available_at"] > arrival_time:
return False
return row["consensus"] is not None and not row["conflict"]
candidate = {"text_revision": "r47", "rule_text_revision": "r47",
"feature_available_at": 12, "consensus": "refund",
"conflict": False}
assert weak_label_row_allowed(candidate, 17)
assert not weak_label_row_allowed({**candidate, "feature_available_at": 19}, 17)
Performance and operating cost
The admission gate is O(1) per row. Comparing every pair of rule outputs for correlation is O(r²) across r functions before sampling; for a modest rule set this can be cheaper than deploying a complex label model. A clean independent holdout costs human review, but it prevents cheap generated labels from producing an expensive false confidence.
Common Mistakes
- Using a post-resolution queue field to predict the intake route.
- Placing copied messages on both sides of the split.
- Assuming several highly correlated votes prove correctness.
- Retiring a rule without tracking which training labels it created.
