An explanation should be checked for repeatability and usefulness before it becomes evidence in a reviewer workflow.
Explanation quality gates: stability, reference data and reviewer use
Test repeatability under a fixed decision
Run the same explanation several times with the same model, feature snapshot, method revision, reference data and seed. Large swings in feature rank or sign may indicate a fragile approximation. Define an acceptable comparison suited to the method; exact byte equality is too strict for some stochastic algorithms, while a vague “looks similar” review is too weak. Provenance is required before stability can be interpreted. A stable output can still be systematically wrong.
Probe realistic neighboring inputs
Perturb only features that can plausibly vary without breaking domain constraints. A merchant-age value cannot be changed independently of the merchant opening date if one is derived from the other. A model may rely on several correlated signals, so an attribution that swaps them in impossible combinations can mislead a reviewer. Test missing fields, scan artifacts and borderline scores. Compare explanation changes with prediction changes, but do not call the comparison a causal intervention.
Evaluate the human task
Ask reviewers whether the explanation helps them find an input error, request missing evidence or decide to override. Measure time to a justified resolution and disagreement against adjudicated outcomes, not just whether they clicked the panel. An explanation can sound persuasive while increasing automation bias. Keep the model score and policy route visible as separate facts, and permit the reviewer to record a reason independent of the explanation. Override records provide a later audit trail.
Treat display changes as releases
A new attribution method, reference population or wording can change reviewer behavior without altering the model. Version the display and test it on frozen decisions, privacy checks and a staged reviewer cohort. Monitor missing explanations, generation failures and unusual override shifts. The project forces an unstable feature ranking and a stale snapshot, then decides whether the panel should display, warn or hold. Operational alerts should identify the owner of a broken explanation service without implying the scorer failed.
Implementation
def explanation_rank_gate(rankings, minimum_top_agreement=0.75):
if not rankings or any(not ranking for ranking in rankings):
raise ValueError("missing rankings")
anchor = rankings[0][0]
agreement = sum(ranking[0] == anchor for ranking in rankings) / len(rankings)
return {"state": "review" if agreement >= minimum_top_agreement
else "hold", "agreement": agreement}
stable = [["scan_quality", "merchant_age"],
["scan_quality", "currency"], ["scan_quality", "merchant_age"],
["scan_quality", "receipt_length"]]
assert explanation_rank_gate(stable)["state"] == "review"
assert explanation_rank_gate(stable[:2] + [["currency"], ["currency"]])[
"state"] == "hold"
Performance and operating cost
Comparing the top feature across r runs is O(r) time and O(1) extra space. Producing those runs can cost r times the explanation computation. Top-rank agreement is only a simple screening diagnostic; it does not establish fidelity, absence of bias or usefulness to reviewers. The release gate needs task-level tests as well.
Common Mistakes
- Assuming a repeated attribution is therefore correct.
- Perturbing correlated features into impossible combinations.
- Testing a panel only for render success and not reviewer decisions.
- Changing reference data without a new explanation revision.
