An explanation must be tied to a specific model, input and method version; similar decisions should not acquire contradictory narratives from harmless perturbations.
Explanation stability and fidelity under model updates
Version the complete explanation path
Store the model checkpoint, feature transform, explanation method, reference data and decision threshold with each generated artifact. A new checkpoint can preserve predictions while changing feature attributions. Comparing explanations without those versions creates a false audit trail. Champion–challenger review offers a model comparison frame.
Measure local fidelity
A surrogate explanation should reproduce the model’s nearby predictions on valid, in-support perturbations. A clean linear chart fitted on impossible combinations is not a faithful account of the model’s actual operating region. The code checks agreement between a declared model and a simple local surrogate on approved temperature settings; it is an arithmetic test, not a general explanation method.
Perturb only benign inputs
Repeat an explanation after small measurement noise, stable image crops or different shuffle seeds that should not change the decision. Track top-feature overlap and whether the recommended action changes. A dramatic change may indicate unstable fitting, correlated features or a decision near the threshold. Grouped perturbation helps interpret one cause.
Do not reward a plausible story
Operators may prefer concise reasons that fit prior beliefs. An explanation audit needs measurable fidelity, feasible actions, consistency and error outcomes, not subjective comfort alone. Keep model error and calibration next to the explanation. Calibration asks whether risk numbers deserve action.
Separate prediction from action evidence
Even a stable and faithful account of what the model uses is not proof that manipulating the named feature improves the real outcome. If the interface proposes a repair, require a separate approved action pathway. Feasibility review is the minimum filter.
Implementation
approved_temperatures = (58, 60, 62, 64)
def deployed_model(temperature):
return 0.013 * temperature - 0.272
def local_surrogate(temperature):
return 0.0131 * temperature - 0.278
def maximum_local_error(settings):
return max(abs(deployed_model(setting) - local_surrogate(setting))
for setting in settings)
error = maximum_local_error(approved_temperatures)
assert round(error, 4) == 0.0004
assert error < 0.01
model_versions = {"checkpoint": "round-11", "features": "v4",
"explanation": "local-linear-v2", "threshold": 0.52}
assert len(model_versions) == 4Performance and operating cost
Evaluating a surrogate on P local points costs O(P) model calls and O(1) storage if errors stream. Repeated explanation runs across records, seeds and model versions can multiply inference cost considerably. A low fidelity error on one tiny range does not establish global reliability or causal meaning.
Common Mistakes
- Do not compare explanation outputs without model and method versions.
- Do not call a persuasive narrative faithful without checking predictions.
- Do not turn a stable attribution into a repair recommendation without action evidence.
