Explainable vision transformer framework for breast lesion classification with Grad-CAM support
摘要
Given that breast cancer remains the most lethal malignancy among women globally, there is an urgent need for reliable and comprehensible diagnostic technology. Deep learning models have demonstrated encouraging outcomes in breast imaging; nevertheless, existing methodologies struggle to encompass global context and deliver clinically relevant justifications for their conclusions.
MethodThe ExViT-Breast framework was trained and internally validated on CBIS-DDSM (n = 1,566 subjects; masses and calcifications), externally validated on INbreast, and evaluated across modalities on the BUSI ultrasound dataset. A Vision Transformer backbone with a cross-view attention fusion module integrated paired craniocaudal (CC) and mediolateral oblique (MLO) mammographic views. Four explainability methods were compared: Grad-CAM, Grad-CAM++, LayerCAM, and Transformer-LRP. Classification performance was assessed by AUC and sensitivity at one false positive per examination. Explanation quality was evaluated using faithfulness (deletion/insertion curves), clinical plausibility (pointing game against expert-defined lesion regions), and ensemble-based uncertainty estimation. Model calibration was quantified using Expected Calibration Error (ECE).
ResultsExViT-Breast achieved superior diagnostic performance with an AUC of 0.94 ± 0.02 on CBIS-DDSM, significantly outperforming conventional CNN architectures (ResNet-50: AUC = 0.89, p < 0.001) and single-view Vision Transformers (AUC = 0.91, p = 0.003). External validation on INbreast demonstrated robust generalization (AUC = 0.91 ± 0.03), while cross-modality evaluation on BUSI yielded AUC = 0.88 ± 0.04. Transformer-LRP consistently provided the highest explanation quality across all datasets, achieving faithfulness scores of 0.71 ± 0.05 and plausibility rates of 79.5% on CBIS-DDSM, significantly superior to Grad-CAM variants (p < 0.001).
ConclusionThe ExViT-Breast framework demonstrates the potential of explainable Vision Transformers in breast imaging, combining high classification performance with clinically interpretable explanations. The integration of multi-view analysis and quantitative explainability assessment positions this approach as a promising tool for computer-aided diagnosis in breast cancer screening and detection.
Clinical trialClinical trial number: not applicable.