Explainable AI in Medical Imaging and Diagnosis: A Critical Review of Methods, Clinical Translation Barriers, and an Integrated Evaluation Framework
DOI:
https://doi.org/10.14740/aicm36Keywords:
Explainable artificial intelligence, Medical imaging, Clinical decision support, Diagnostic AI, Interpretability, Trust calibration, Radiology, Digital pathologyAbstract
Deep learning now matches or outperforms specialist clinicians on many medical imaging benchmarks, but clinical uptake has lagged well behind the benchmark numbers. Opacity is usually blamed. Explainable artificial intelligence (XAI) is the field’s proposed fix, and it has produced a large and still-growing toolkit: saliency maps, feature-attribution scores, concept-based explainers, and counterfactual image generators, spanning radiology, digital pathology, dermatology, and ophthalmology. This critical review works through more than 35 recent primary studies, systematic reviews, and meta-analyses to test an assumption the literature rarely states outright but almost always relies on: that producing an explanation is the same thing as producing understanding. It is not, at least not reliably. Quantitative fidelity and stability data, human-factors evidence on automation bias and trust calibration, and the regulatory direction set by the US Food and Drug Administration and the EU Artificial Intelligence Act all point the same way: widely used post-hoc methods such as Grad-CAM, LIME, and SHAP routinely fail basic faithfulness checks, explanations can make clinicians more confident in wrong predictions rather than less, and the field currently evaluates all this along technical, psychological, and regulatory tracks that barely intersect. The review’s central contribution is an integrated Clinical-XAI Evaluation Framework built around six dimensions, fidelity, robustness, plausibility, actionability, equity, and regulatory readiness, meant to give editors, developers, and clinical adopters one shared checklist for deciding whether an explanation has actually earned a place in a diagnostic workflow, rather than several separate and often contradictory ones. The review argues that the field should shift its default investment away from post-hoc explainability and toward inherently interpretable architectures, modality-specific evaluation protocols, and prospective validation tied to decision outcomes, not further proliferation of attribution maps that look convincing but have rarely been checked against what the model actually did.
Published
Issue
Section
License
Copyright (c) 2026 The authors

This work is licensed under a Creative Commons Attribution 4.0 International License.





