Reconstruct one decision, reject stale explanation evidence and test whether reviewers can use the corrected panel without over-trusting it.
Project: audit an explanation for a disputed receipt decision
Freeze the disputed decision
A receipt was routed to review by a fixed risk model and threshold policy. Preserve decision ID, score, route, model digest, feature snapshot, preprocessor, policy revision and observation time. Reconstruct only from permitted retained data. The explanation service must use the same identity and its own method revision and reference set. The provenance gate rejects a panel generated from a newer merchant feature snapshot, even if the resulting feature list looks plausible.
Challenge the explanation method
Generate repeated explanations under a fixed seed policy and inspect top-feature agreement. A scanner-quality signal is stable, but merchant age and receipt length trade places across runs. Do not present a single run as a definitive cause. Create a nearby input with a physically plausible scan-quality change and compare predictions; avoid impossible combinations of derived fields. Stability tests support a warning or hold, not a claim of causal effect.
Exercise reviewer decisions
Give reviewers the original receipt evidence, proposed route and a panel marked with its scope and uncertainty. Ask them to identify whether missing scan quality needs re-capture, whether an override is justified and what evidence they used. Keep their decision separate from later adjudication. The review ledger records an override reason but does not turn it into an outcome label. Compare panel-on and panel-off review of similar cases without treating fewer clicks as proof of better decisions.
Deliver a release disposition
Publish provenance checks, retained snapshot coverage, method revision, stability result, privacy review, reviewer-task outcome and known limits. The stale-snapshot explanation is rejected; the corrected panel may proceed only with a warning if its unstable ranking is material to the task. Record a rollback of the panel separately from the model and score policy. Link future cases to the model scope review so explanations never extend the approved use of the model.
Implementation
def explanation_audit(decision, panel, top_feature_agreement):
if panel["model_digest"] != decision["model_digest"]:
return "reject:model"
if panel["feature_snapshot"] != decision["feature_snapshot"]:
return "reject:snapshot"
if not 0 <= top_feature_agreement <= 1:
raise ValueError("invalid agreement")
if top_feature_agreement < 0.75:
return "hold:unstable"
return "review-panel"
decision = {"model_digest": "risk-47", "feature_snapshot": "snap-82"}
panel = {**decision, "method_revision": "local-r4"}
assert explanation_audit(decision, panel, 0.82) == "review-panel"
assert explanation_audit(decision, {**panel,
"feature_snapshot": "snap-83"}, 0.82) == "reject:snapshot"
assert explanation_audit(decision, panel, 0.5) == "hold:unstable"
Performance and operating cost
The audit gate is O(1) time and space. The complete exercise costs repeated explanation runs, retained input storage, privacy review and reviewer time. A fixed rank threshold is a screening rule, not a measure of explanation truth; the final disposition also depends on method validity and whether the panel helps the actual review task.
Common Mistakes
- Explaining a disputed decision from a later feature snapshot.
- Describing a feature contribution as a causal reason.
- Using reviewer agreement alone as a fidelity metric.
- Rolling back the model when only the explanation panel is unsafe.
Read next
- Production explanations: pin the exact decision and method
- Explanation quality gates: stability, reference data and reviewer use
- Human overrides: keep decisions, reasons and labels distinct
- Project: approve a receipt model with a scoped evidence dossier
- Inference logs: keep diagnostic joins without copying sensitive payloads
