AI in Clinical Medicine, ISSN 2819-7437 online, Open Access
Article copyright, the authors; Journal compilation copyright, AI Clin Med and Elmer Press Inc
Journal website https://aicm.elmerpub.com

Review

Volume 2, August 2026, e36


Explainable AI in Medical Imaging and Diagnosis: A Critical Review of Methods, Clinical Translation Barriers, and an Integrated Evaluation Framework

Figures

↓  Figure 1. Taxonomy of explainable AI (XAI) methods applied to medical imaging, organized by the intrinsic/post-hoc and model-specific/model-agnostic axes described in Section “A Taxonomy of Explainability in Medical Imaging AI.” Italicized notes summarize the fidelity evidence discussed in the text.
Figure 1.
↓  Figure 2. Fidelity versus plausibility: a conceptual illustration of the review’s central argument. Positions are qualitative, based on the evidence discussed in the text, not derived from a single quantitative scale.
Figure 2.
↓  Figure 3. Perturbation-based fidelity scores (mean, 95% CI) for LIME, Grad-CAM, and SHAP, pooled across 67 studies spanning radiology, pathology, and ophthalmology applications. Source: meta-analytic evidence discussed in Section “The Fidelity Problem: Do Explanations Explain?” [1].
Figure 3.
↓  Figure 4. The six-dimension Clinical-XAI Evaluation Framework proposed in Section “Toward an Integrated Clinical-XAI Evaluation Framework”. Each dimension is assessed independently rather than combined into a single composite score; see Section “Applying the framework: a worked example” for a worked example applying all six to two contrasting cases.
Figure 4.

Tables

↓  Table 1. Modality-Specific Summary of Dominant XAI Approach, Strongest Available Evidence, and Key Limitation
 
ModalityDominant XAI approachStrongest evidence identifiedKey limitation/gap
RadiologySaliency-based (Grad-CAM); adoption of any explainability analysis remains limited overallProfessional society (RSNA) advocacy for stronger transparency and post-deployment monitoring requirements.A substantial share of published radiology AI studies report no explainability analysis at all [26, 27].
Digital pathologyCounterfactual tissue-region perturbation (HIPPO)HIPPO validated across breast metastasis detection, melanoma/breast cancer prognostication, and glioma mutation classification [22].Patch-based tiling breaks biological continuity; explainability must be paired with causability, not treated as equivalent to it [29].
DermatologyConcept-based methods; text- and region-based explanations116-dermatologist, three-phase reader study: explainable AI significantly raised diagnostic confidence and trust versus an unexplained system [30].Evidence base still leans on non-clinical dermoscopic imagery, with limited attention to skin-tone representation.
OphthalmologyInherently interpretable architectures (sparse BagNet, ExplAIn)Matched black-box accuracy with no post-hoc stage; highest-precision native evidence maps of any modality reviewed [24, 25].Post-hoc methods show the worst noise-stability degradation of any modality (53% for SHAP) [1].

 

↓  Table 2. Regulatory Posture Toward Explainability in Diagnostic Imaging AI Across Three Jurisdictions
 
JurisdictionGoverning frameworkRisk classification for diagnostic imaging AIExplainability requirement
European UnionEU Artificial Intelligence Act“High-risk”Conformity assessment before market approval; mandated human-in-the-loop oversight; post-market monitoring with dataset traceability; restrictions on fully automated decisions without human oversight.
United StatesFDA review process for AI/ML-based Software as a Medical DeviceCase-by-case clearanceNo single explicit explainability mandate; professional societies (e.g., RSNA) have formally urged stronger transparency, explainability, and lifecycle-management guidance.
Gulf Cooperation Council/MENANational digital-health and AI strategies (varies by country)Not yet formally codified for diagnostic imaging AI specificallyNot yet codified; regulatory infrastructure is, in most cases, still developing relative to clinical AI investment.

 

↓  Table 3. Operational Definitions and Example Measurable Evaluation Criteria for the six Clinical-XAI Evaluation Framework Dimensions
 
DimensionOperational definitionExample measurable criterionEvidence this dimension draws on
1. FidelityDegree to which the explanation reflects the model’s actual computation, not merely a plausible-looking approximation of itPasses a weight/label-randomization sanity check [15]; perturbation-based faithfulness score reported with a stated threshold (e.g., \u22650.6) [1]Section “The Fidelity Problem: Do Explanations Explain?”
2. Robustness & stabilityConsistency of the explanation under small, clinically irrelevant input perturbations and repeated stochastic runs\u226420% change in explained region under \u226410% input noise, tested per modality rather than assumed from another task [1]Section “The Fidelity Problem: Do Explanations Explain?”
3. PlausibilityAlignment of the explanation with clinically recognized reasoning, judged independently of fidelityBlinded expert-rater agreement (e.g., Cohen's \u03ba \u22650.6) that the highlighted evidence matches recognized diagnostic criteria [30]Section “Human Factors: Trust, Automation Bias, and the Limits of Explanation”
4. ActionabilityWhether the explanation measurably changes clinician behavior or downstream outcomes, not only stated confidencePre-registered outcome (e.g., reading time, diagnostic accuracy, or referral appropriateness) shown to improve under AI support with the explanation present [24]Section “Human Factors: Trust, Automation Bias, and the Limits of Explanation”
5. EquityConsistency of both model performance and explanation quality across patient subgroups and acquisition equipmentSubgroup performance gap (e.g., across skin tone, sex, age, scanner type) reported and below a pre-specified threshold, not merely aggregate accuracySection “Cross-Cutting Barriers”
6. Regulatory readinessWhether the explanation methodology and its evidence would satisfy applicable jurisdictional documentation and oversight requirementsExplicit mapping of the evidence above to EU AI Act Article-level conformity requirements and/or FDA SaMD guidance, including a stated plan for re-validation after model updatesSection “Regulatory and Governance Landscape”

 

↓  Table 4. The Six-Dimension Clinical-XAI Evaluation Framework Applied to Two Contrasting Cases, as Worked Through in Section “Applying the Framework: A Worked Example”
 
DimensionSparse BagNet for diabetic retinopathy screening [24]Generic Grad-CAM for chest radiograph triage
1. FidelitySatisfied by construction—the evidence map is the model’s actual computation, not an approximationFails by default without a study-specific fidelity check; meta-analytic base rate is only 0.54 [1]
2. Robustness & stabilityPlausible given the architecture, but never systematically stress-tested against noiseDoubtful—fidelity documented to degrade sharply on transformer backbones [7]
3. PlausibilityStrong—evidence regions matched ophthalmologist-identified lesions with high precisionTends to be high, but this is exactly what makes an unchecked fidelity gap dangerous
4. ActionabilityWell satisfied—screening speed and accuracy under AI support measured directly [24]Rarely measured in studies that report only qualitative heatmaps
5. EquityThin—static, single-institution trainingRarely assessed in typical Grad-CAM studies
6. Regulatory readinessUncertain—unclear how evidence maps hold up under continuous post-market monitoringUncertain—no fidelity documentation exists to build a regulatory case on

 

↓  Table 5. Comparative Overview of Major XAI Method Families Discussed in This Review
 
Method familyRepresentative techniquesFidelity evidencePrimary clinical risk
Gradient/saliency-based (post-hoc, model-specific)Grad-CAM, Grad-CAM++, guided backpropagation, Integrated GradientsModerate and architecture-dependent; fidelity ∼0.54 for Grad-CAM, degrades further on transformer backbones [1, 7]Plausible-looking but unfaithful heatmaps may foster false confidence
Perturbation-based (post-hoc, model-agnostic)LIME, SHAPVariable: LIME ∼0.81, SHAP ∼0.38 in meta-analytic evidence; both unstable under noise, especially SHAP in ophthalmology [1]Explanation instability across repeated runs undermines medico-legal defensibility
Concept-basedConcept Bottleneck Models, concept activation vectors, DermXHigh plausibility where concept annotations are clinically grounded; fidelity bounded by concept-dataset quality [17, 18, 21]Concept vocabulary incompleteness or annotation bias limits generalizability
Counterfactual/example-basedHIPPO tissue counterfactuals, retinal image counterfactual generation, case-retrieval explainersStrong for model auditing and bias detection; limited large-scale clinical reader validation to date [22, 23]Synthetic counterfactual images may not represent clinically realistic disease presentations
Inherently interpretable architecturesSparse BagNet, ExplAIn lesion-segmentation modelExplanation is definitionally faithful to computation; strongest available reader-study evidence of clinical benefit [24, 25]Architecture constraints may limit applicability to tasks beyond well-engineered use cases