Weak supervision assigns provisional labels through explicit rules that may vote, conflict or abstain; coverage and agreement are diagnostics, not proof of label quality.
Weak-label rules, coverage and conflicts
Write rules from observable evidence
A scan gap may indicate a missed handoff, while ample crew may indicate an on-time case. A rule should inspect only fields available at the intended labeling point and return positive, negative or abstain. The code creates three small rules and measures how often any rule votes. It does not convert silence into a negative label. Feature timing prevents a future outcome field from quietly becoming a rule input.
Inspect conflicts before combining votes
Two rules can disagree on the same shipment. That can reveal different operational meanings, inconsistent feature clocks or simply weak heuristics. A majority vote is one combination policy, but it ignores rule accuracy and correlation. The code leaves ties unassigned. Several rules copied from the same scan system do not provide independent confirmation. Keep a matrix of individual votes and a conflict count before generating any training target.
Audit against genuinely reviewed cases
Choose a stratified, time- and site-aware set of human-reviewed shipments. Estimate each rule’s positive and negative error, coverage and drift across sites; do not tune rules repeatedly on the final test labels. The weak labels may help train a candidate, but evaluation must use independent genuine outcomes. Rare-event metrics explain why a rule with high overall agreement may still miss most failures.
Keep rule provenance and ownership
Store the rule version, fields used, execution time and reason for each vote. When the scan pipeline changes, a once-useful rule can degrade without an exception. Review rule outputs after ingestion changes and retain the original human label where one exists. The monitoring guide separates early input alarms from delayed outcome evidence.
Compare with other label-budget tactics
Pseudo-labels come from a model’s own predictions; active learning sends selected cases to a reviewer; weak rules use domain logic. These approaches can be combined, but their labels have different trust levels and costs. Do not merge all three into one column without provenance. Self-training and active review are distinct workflows.
Implementation
# Each rule returns 1, 0, or None for abstention.
shipments = [
{"id": "S301", "scan_gap": 4, "crew": 8},
{"id": "S302", "scan_gap": 0, "crew": 9},
{"id": "S303", "scan_gap": 3, "crew": 2},
{"id": "S304", "scan_gap": 1, "crew": 5},
]
def gap_rule(row):
return 1 if row["scan_gap"] >= 3 else None
def low_crew_rule(row):
return 1 if row["crew"] <= 3 else None
def well_staffed_rule(row):
return 0 if row["crew"] >= 7 else None
rules = (gap_rule, low_crew_rule, well_staffed_rule)
votes = {row["id"]: [rule(row) for rule in rules] for row in shipments}
covered = sum(any(vote is not None for vote in row_votes)
for row_votes in votes.values())
conflicts = sum(len({vote for vote in row_votes if vote is not None}) > 1
for row_votes in votes.values())
def majority_or_abstain(row_votes):
positives = row_votes.count(1)
negatives = row_votes.count(0)
return 1 if positives > negatives else 0 if negatives > positives else None
provisional = {shipment_id: majority_or_abstain(row_votes)
for shipment_id, row_votes in votes.items()}
assert covered == 3
assert conflicts == 1
assert provisional["S301"] is None
assert provisional["S304"] is NonePerformance and operating cost
Applying R rules to N rows costs O(NR) time and O(NR) storage if every vote is retained. A final majority is cheap, but gold-label audits, conflict analysis and rule maintenance dominate the real operating cost. Related rules can make naive vote counts misleading.
Common Mistakes
- Do not treat an abstention as a negative label.
- Do not count correlated rules as independent evidence.
- Do not evaluate weak labels against the same reviewed cases used to tune the rules.
Read next
- Pseudo-label selection and contamination control
- Active learning with uncertainty and diversity
- Feature drift and delayed-label monitoring
- Label budget and incremental learning project
Continue the workflow: Positive–unlabeled target and observation state.
