Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Leakage-safe preprocessing: fit every learned transform inside the training fold

Last updated: 5 Oct 20265 min read
tutorial
BeginnerBy AITrove Editorial

A preprocessing Pipeline learns imputation, scaling and encoding from training rows and applies the same fitted transformation to validation and serving rows.

Separate fit from transform

A median imputer learns a value from observed data. If it sees the entire dataset before train/test splitting, the held-out rows influence that value. Feature selection, vocabulary building and dimensionality reduction have the same risk. Define the split first. Fit a Pipeline on each training fold; use its predict or transform methods on held-out rows without a second fit. The split strategy] still needs to respect stores and time.

Make columns explicit

A receipt model has numeric amount and account-age fields plus a categorical intake channel. Use a ColumnTransformer that names those columns and gives each type its own imputation and encoding contract. Ignore an unknown category at serving only if the business accepts that fallback; count such cases. Keep a frozen list of expected columns so a renamed field fails early rather than silently shifting feature positions.

Keep the whole artifact together

Persist the fitted preprocessing steps and estimator as one versioned artifact. A model serialized without its encoder cannot interpret live values consistently. Save the training schema, dependency versions and validation split with the artifact. Only load artifacts from trusted locations; generic Python object deserialization is not a safe exchange format for untrusted files.

Prove boundary behavior

Train on a fold where one category is absent, then score a validation row containing that category. Check the intended behavior and count. Add a row with a missing amount and confirm the imputer uses the training median. The missing-value policy] still needs a domain justification; the Pipeline only enforces it consistently.

Implementation

python
numeric = Pipeline([("impute", SimpleImputer(strategy="median")),
                    ("scale", StandardScaler())])
category = Pipeline([("impute", SimpleImputer(strategy="most_frequent")),
                     ("encode", OneHotEncoder(handle_unknown="ignore"))])
prep = ColumnTransformer([
    ("numeric", numeric, ["receipt_amount", "account_age_days"]),
    ("channel", category, ["intake_channel"]),
])
model = Pipeline([("prepare", prep),
                  ("classify", LogisticRegression(max_iter=800))])
model.fit(train_rows, train_labels)
validation_scores = model.predict_proba(validation_rows)[:, 1]

Performance and operating cost

Fitting scales roughly with rows, columns and estimator choice; one-hot encoding can expand a high-cardinality category into many sparse columns. Cross-validation repeats the full fit for each fold and keeps transformed matrices in memory.

Common Mistakes

  • Do not fit preprocessing on the full dataset before splitting.
  • Do not save only the classifier while recreating the encoder by hand.
  • Do not treat unknown categories as harmless without measuring their rate.

Read next

Connected implementation

Continue the workflow: Sparse text baselines: fit vocabulary inside the training boundary.

Continue the workflow: Train-only imputation and missingness indicators.

Continue the workflow: Ridge regularization with training-only scaling.

Continue the workflow: Distance scaling before clustering.

Continue the workflow: Fold-local feature selection and stability.

machine-learning
leakage-safe-preprocessing-pipeline
Storage details