Feature selection ranks or tests candidate inputs using training data inside each validation fold, then checks whether the chosen set and gain survive other folds.
Fold-local feature selection and stability
Treat a selector as fitted state
Suppose a handoff model considers backlog, crew count and route congestion. A selector that compares each feature with the missed-handoff label has learned from those labels, even if no classifier was fitted yet. Running it on the whole dataset before cross-validation leaks validation outcomes into every training fold. Fit the selector inside each fold, then transform that fold’s validation rows using only its selected columns. The pipeline guide applies to selectors too.
Use a score that matches its claim
The code ranks numeric features by the absolute difference between positive and negative training means. It is a univariate screening example, not a feature-importance estimate, proof of causality or optimal subset search. Correlated features can split value, interactions can be missed, and a rare class can make the mean unstable. The classifier’s actual validation score, cost and slice performance decide whether a selected set is useful.
Measure selection stability
Record selected names across time-aware or group-aware development folds. If one feature wins only when a single warehouse enters training, investigate measurement and site dependence. A stable shortlist can reduce maintenance and inference cost, but stability alone does not guarantee better predictions. Held-out permutation importance asks a different question about a fitted model.
Reject future-only predictors first
No statistical selector can make an unavailable field deployable. Remove outcomes, post-handoff scan counts and identifiers that encode the label before ranking. For each remaining feature, record capture time, missingness and feature version. A strong score from a leaking field is evidence of a broken experiment. The clock check precedes feature search.
Freeze the final feature contract
After choosing the selection method and its settings on development folds, refit it on permitted development rows and evaluate the complete pipeline once on a sealed future test. Keep the chosen feature order with the trained model. Re-selecting columns from live outcomes without a new validation process changes the model. Untouched-test selection defines that final step.
Implementation
from collections import Counter
development_rows = [
{"backlog": 8, "crew": 8, "congestion": 2, "missed": 0},
{"backlog": 12, "crew": 7, "congestion": 3, "missed": 0},
{"backlog": 17, "crew": 6, "congestion": 4, "missed": 0},
{"backlog": 22, "crew": 5, "congestion": 4, "missed": 1},
{"backlog": 27, "crew": 4, "congestion": 6, "missed": 1},
{"backlog": 31, "crew": 3, "congestion": 7, "missed": 1},
]
candidate_names = ("backlog", "crew", "congestion")
def choose_feature(training_rows):
positives = [row for row in training_rows if row["missed"] == 1]
negatives = [row for row in training_rows if row["missed"] == 0]
if not positives or not negatives:
raise ValueError("both classes required")
gaps = {name: abs(sum(row[name] for row in positives) / len(positives)
- sum(row[name] for row in negatives) / len(negatives))
for name in candidate_names}
return max(candidate_names, key=lambda name: (gaps[name], name))
folds = [([0, 1, 3, 4], [2, 5]), ([0, 2, 3, 5], [1, 4])]
selected = []
for training_ids, validation_ids in folds:
assert set(training_ids).isdisjoint(validation_ids)
selected.append(choose_feature([development_rows[index] for index in training_ids]))
selection_counts = Counter(selected)
assert len(selected) == 2
assert set(selection_counts).issubset(candidate_names)
assert all(count > 0 for count in selection_counts.values())Performance and operating cost
This univariate screen takes O(KNF) time across K folds, N rows per fold and F features, with O(F) score memory. More complex subset search can be much costlier. The larger expense may be repeated downstream model fitting, so budget selection and evaluation together.
Common Mistakes
- Do not select features once on all labeled rows before cross-validation.
- Do not keep a future-only field because its ranking score is high.
- Do not treat a selected feature as a causal driver.
