A validation split should test generalization to the actual next entity or future period, rather than mix dependent observations across partitions.
Group and time validation: split by the failure you expect in production
Match the deployment question
If a receipt model will serve stores absent from training, hold out entire stores. If it will score next month’s submissions at known stores, hold out later time. A random row split can place near-identical receipts from one store on both sides and give an optimistic estimate. GroupKFold prevents a group from appearing in both sides of a fold; TimeSeriesSplit preserves order for equally spaced samples, but it does not understand arbitrary event-time delays or group ownership automatically.
Fence labels and transforms
Set a cutoff for feature availability and label maturity. A receipt submitted just before the fold boundary may receive its review outcome after the validation period begins; impose a gap or embargo where needed. Fit preprocessing only on each training fold. A Pipeline] applies learned transformations inside each fold. If there are fewer unique groups than folds, revise the plan rather than forcing a split.
Keep one final answer unseen
Use development folds to compare features and tune hyperparameters. After selection, evaluate once on a separately held-out future period or group set. Do not keep changing the model after seeing that final score without treating the holdout as consumed. Report fold sizes, group counts, positive rates and the worst important slice; an average alone can conceal failure in a small region.
Verify the index sets
Before fitting, assert no store ID occurs in both train and validation for a group split, and assert every validation timestamp follows the training window for a chronological split. Save these indices with the run. Cohort definitions] should align the evaluation population with the population the service will actually see.
Implementation
from sklearn.base import clone
from sklearn.model_selection import GroupKFold
splitter = GroupKFold(n_splits=4)
for train_index, validation_index in splitter.split(
features, labels, groups=receipts["store_id"]):
train_stores = set(receipts.iloc[train_index]["store_id"])
validation_stores = set(receipts.iloc[validation_index]["store_id"])
assert train_stores.isdisjoint(validation_stores)
fold_model = clone(model_template)
fold_model.fit(features.iloc[train_index], labels.iloc[train_index])
predictions = fold_model.predict(features.iloc[validation_index])Performance and operating cost
K folds multiply model-fit work by roughly K and require saving or recomputing each prediction. Grouping can produce uneven fold sizes, so compare row and group counts rather than assuming identical test mass.
Common Mistakes
- Do not random-split dependent records from one entity across train and test.
- Do not fit an imputer or encoder before cross-validation.
- Do not reuse the final holdout as a tuning dashboard.
Read next
- Prediction-time feature availability: reject future information before training
- Leakage-safe preprocessing: fit every learned transform inside the training fold
- Decision thresholds: choose an action from probabilities and error costs
- Bootstrap intervals: estimate uncertainty at the right sampling unit
Connected implementation
Continue the workflow: Image augmentation: split originals first and preserve the label.
Continue the workflow: Text validation: split conversations, duplicates and time together.
Continue the workflow: Experiment design: assign the right unit and guard against interference.
Continue the workflow: Rolling-origin backtests: rehearse the forecast as it would have run.
Continue the workflow: Vision evaluation splits: group captures and test acquisition shift.
Continue the workflow: Geospatial evaluation: spatial holdouts, leakage and transfer.
Continue the workflow: Adjudication and gold sets: versioned reference decisions.
Continue the workflow: Decision-tree splits and minimum leaf support.
Continue the workflow: Rolling-origin retraining with an outcome embargo.
Continue the workflow: Sequence-label contracts and token alignment.
Continue the workflow: Survival risk sets and landmark snapshots.
Continue the workflow: Active-learning evaluation and stopping by value.
Continue the workflow: S-learners and T-learners for treatment-effect baselines.
