Permutation importance measures how much a fitted model’s held-out error changes when one feature column is shuffled while the model remains fixed.
Permutation importance on held-out data
Start with a useful model and valid data
A clearance-time model predicts from backlog and route distance. First confirm that it beats a training-only baseline on a future holdout; feature importance for a poor model can be misleading. Preserve the model and holdout rows. For one feature, shuffle values across holdout records, predict again, and compare loss. The model is not retrained. The baseline lesson establishes the reference error.
Use repeats rather than one lucky shuffle
One permutation can accidentally leave values close to their original locations, especially in a short holdout. Repeat with fixed, distinct random seeds and report the distribution of loss increases. The code uses three repeats for a compact demonstration and measures MAE increase; a real analysis needs enough records and repetitions for stable comparison. A negative increase means the shuffled feature happened to improve this measured loss, not that the feature is harmful in every setting.
Keep the interpretation predictive
A large importance means the fitted model relied on that feature for this holdout distribution and metric. It does not prove the feature causes longer clearance times. Correlated features can substitute for each other, making each individual permutation appear weak even when the pair carries useful signal. The shuffle can also produce feature combinations that rarely occur in real shipments. Tree rules have related interpretation limits.
Respect entities and time
If rows share a route or customer, shuffling across unrelated groups may create impossible profiles. One may permute within declared strata, but then the question changes to importance conditional on that grouping. Do not use test data repeatedly to select features; reserve a development holdout for diagnostic selection and keep a separate final test. Split design still governs inspection.
Use the result to test a hypothesis
If a feature thought essential has near-zero importance, check whether it is redundant, broken at serving time or suppressed by preprocessing. If a future-only feature appears dominant, remove it before celebrating. Document feature version, loss, seeds and holdout support. The review project asks for a concrete follow-up rather than a decorative ranking.
Implementation
from random import Random
def mean_absolute_error(actual, predicted):
if not actual or len(actual) != len(predicted):
raise ValueError("aligned outcomes required")
return sum(abs(left - right) for left, right in zip(actual, predicted)) / len(actual)
def backlog_predictor(shipment):
return shipment["backlog"] / 4
def backlog_permutation_increase(holdout_shipments, seeds):
actual = [shipment["clearance_hours"] for shipment in holdout_shipments]
baseline = mean_absolute_error(actual, [backlog_predictor(shipment)
for shipment in holdout_shipments])
changes = []
for seed in seeds:
values = [shipment["backlog"] for shipment in holdout_shipments]
Random(seed).shuffle(values)
altered = [dict(shipment, backlog=value)
for shipment, value in zip(holdout_shipments, values)]
changed_error = mean_absolute_error(actual, [backlog_predictor(shipment)
for shipment in altered])
changes.append(changed_error - baseline)
return changes
shipments = [{"backlog": hours * 4, "clearance_hours": hours}
for hours in (3, 5, 8, 12, 15)]
increases = backlog_permutation_increase(shipments, (23, 47, 61))
assert len(increases) == 3 and all(increase >= 0 for increase in increases)
assert any(increase > 0 for increase in increases)Performance and operating cost
With R repeats and N holdout rows, direct permutation evaluation costs O(R × N) time and O(N + R) space. Each additional feature multiplies prediction calls; model serving time can dominate the shuffle itself.
Common Mistakes
- Do not present permutation importance as a causal effect.
- Do not compute importance before verifying holdout predictive quality.
- Do not shuffle across operational groups without considering impossible combinations.
Read next
- Hyperparameter search and an untouched final test
- Prediction-time feature availability: reject future information before training
- Decision-tree splits and minimum leaf support
- Project: select a model and audit warehouse segments
Continue the workflow: Correlated feature permutation audit.
