Permutation importance measures a fitted model’s score change when one input is disrupted; correlated inputs can mask reliance or create implausible records.
Correlated feature permutation audit
Name the inspected quantity
A pump follow-up model uses vibration amplitude and a derived vibration severity band. If both encode nearly the same information, shuffling amplitude alone may barely change accuracy because severity remains. A small single-feature importance is not evidence that vibration is irrelevant. Held-out permutation importance supplies the baseline procedure.
Measure groups before individual fields
For strongly related fields, permute the group together using the same row mapping. This asks how much predictive performance depends on their combined information. Repeat mappings and show the spread of score changes; one small shuffled sample can give an unstable rank. Use validation rows that were not used to fit the model or choose its settings.
Respect real combinations
An unrestricted shuffle can pair high vibration with a severity band that means low vibration. The model was never intended to score such combinations. Grouped permutation preserves the two-field relation within donor rows, but it can still pair them with incompatible operating conditions. A conditional shuffle within equipment family may reduce that issue while changing the question to within-family importance.
Separate reliance from cause
If the model score drops after shuffling a field, the model used information correlated with that field. It does not show that changing the physical pump reading would change failure risk or that technicians should alter that reading. Causal timing is a different analysis.
Compare policy-relevant slices
Run the audit by depot, equipment age and sensor revision. A field can be useful overall yet act as a brittle site marker. Pair explanation changes with held-out error and missingness checks; a compelling chart cannot rescue a model that misses failures. Stability checks bound the claim.
Implementation
inspection_rows = [
{"vibration": 2.2, "severity": "low", "equipment": "pump-A"},
{"vibration": 6.8, "severity": "high", "equipment": "pump-A"},
{"vibration": 3.1, "severity": "low", "equipment": "pump-B"},
{"vibration": 7.4, "severity": "high", "equipment": "pump-B"},
]
def grouped_swap_within_equipment(rows):
by_equipment = {}
for row in rows:
by_equipment.setdefault(row["equipment"], []).append(row)
changed = []
for group in by_equipment.values():
donors = list(reversed(group))
for recipient, donor in zip(group, donors):
changed.append({**recipient, "vibration": donor["vibration"],
"severity": donor["severity"]})
return changed
changed_rows = grouped_swap_within_equipment(inspection_rows)
assert len(changed_rows) == 4
assert all((row["vibration"] >= 5) == (row["severity"] == "high") for row in changed_rows)
assert changed_rows[0]["equipment"] == "pump-A"Performance and operating cost
One grouped shuffle is O(N) expected time and O(N) memory for N rows. Repeating it R times and scoring a model costs O(R times model inference over N rows). The example preserves two-field consistency only; a full conditional permutation needs a defensible conditioning scheme and enough support in every stratum.
Common Mistakes
- Do not interpret a low individual score as proof a correlated signal is unused.
- Do not claim the shuffle estimates a causal effect.
- Do not hide synthetic, out-of-support combinations created by the perturbation.
