A preprocessing Pipeline learns imputation, scaling and encoding from training rows and applies the same fitted transformation to validation and serving rows.
Leakage-safe preprocessing: fit every learned transform inside the training fold
Separate fit from transform
A median imputer learns a value from observed data. If it sees the entire dataset before train/test splitting, the held-out rows influence that value. Feature selection, vocabulary building and dimensionality reduction have the same risk. Define the split first. Fit a Pipeline on each training fold; use its predict or transform methods on held-out rows without a second fit. The split strategy] still needs to respect stores and time.
Make columns explicit
A receipt model has numeric amount and account-age fields plus a categorical intake channel. Use a ColumnTransformer that names those columns and gives each type its own imputation and encoding contract. Ignore an unknown category at serving only if the business accepts that fallback; count such cases. Keep a frozen list of expected columns so a renamed field fails early rather than silently shifting feature positions.
Keep the whole artifact together
Persist the fitted preprocessing steps and estimator as one versioned artifact. A model serialized without its encoder cannot interpret live values consistently. Save the training schema, dependency versions and validation split with the artifact. Only load artifacts from trusted locations; generic Python object deserialization is not a safe exchange format for untrusted files.
Prove boundary behavior
Train on a fold where one category is absent, then score a validation row containing that category. Check the intended behavior and count. Add a row with a missing amount and confirm the imputer uses the training median. The missing-value policy] still needs a domain justification; the Pipeline only enforces it consistently.
Implementation
numeric = Pipeline([("impute", SimpleImputer(strategy="median")),
("scale", StandardScaler())])
category = Pipeline([("impute", SimpleImputer(strategy="most_frequent")),
("encode", OneHotEncoder(handle_unknown="ignore"))])
prep = ColumnTransformer([
("numeric", numeric, ["receipt_amount", "account_age_days"]),
("channel", category, ["intake_channel"]),
])
model = Pipeline([("prepare", prep),
("classify", LogisticRegression(max_iter=800))])
model.fit(train_rows, train_labels)
validation_scores = model.predict_proba(validation_rows)[:, 1]Performance and operating cost
Fitting scales roughly with rows, columns and estimator choice; one-hot encoding can expand a high-cardinality category into many sparse columns. Cross-validation repeats the full fit for each fold and keeps transformed matrices in memory.
Common Mistakes
- Do not fit preprocessing on the full dataset before splitting.
- Do not save only the classifier while recreating the encoder by hand.
- Do not treat unknown categories as harmless without measuring their rate.
Read next
- Group and time validation: split by the failure you expect in production
- Prediction-time feature availability: reject future information before training
- Probability calibration: test whether risk scores mean what they say
- Missing data policy: distinguish absence from a measured zero
Connected implementation
Continue the workflow: Sparse text baselines: fit vocabulary inside the training boundary.
Continue the workflow: Train-only imputation and missingness indicators.
Continue the workflow: Ridge regularization with training-only scaling.
Continue the workflow: Distance scaling before clustering.
Continue the workflow: Fold-local feature selection and stability.
