<p>This paper introduces Grad-TransUNet, a novel deep neural network model designed to improve medical image segmentation by integrating transformers and Grad-CAM with traditional CNN architectures. The study addresses the limitations of conventional CNNs in capturing long-range dependencies and global context in medical images. By incorporating transformers within a hybrid CNN-Transformer encoder, the model leverages multi-head self-attention and multi-layer perceptron blocks to enhance segmentation performance. Additionally, the Grad-CAM technique is employed to generate informative heatmaps, enhancing the model’s explainability. The proposed model fuses the local features of CNN model by explainable features of Grad-CAM and then feeds these enhanced local futures to a transformer in order to enrich global features. Finally, a Cascaded Upsampler (CUP) decoder for transforming encoded features to extract segmentation masks. Extensive experiments on the Kvasir and brain tumor MRI datasets demonstrate significant performance improvements over the state-of-the-art models across metrics such as the dice similarity coefficient (DSC), mean Intersection over Union (mIoU), precision, and recall. In The Kvasir Dataset, the DSC increased by 2.6%, while the IOU exhibited the most significant growth at 6%. Additionally, Recall saw an enhancement of 4.1%, and Precision improved by 3.4%. In the brain tumor MRI dataset, similar advancements were observed. The DSC rose by about 2%, and the IOU grew by 3.2%. Furthermore, Recall improved by 3%, and Precision showed a 1.1% increase.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Explainable AI enhanced transformer based UNet for medical images segmentation using gradient weighted class activation map

  • Abtin Hasannezhad,
  • Saeed Sharifian

摘要

This paper introduces Grad-TransUNet, a novel deep neural network model designed to improve medical image segmentation by integrating transformers and Grad-CAM with traditional CNN architectures. The study addresses the limitations of conventional CNNs in capturing long-range dependencies and global context in medical images. By incorporating transformers within a hybrid CNN-Transformer encoder, the model leverages multi-head self-attention and multi-layer perceptron blocks to enhance segmentation performance. Additionally, the Grad-CAM technique is employed to generate informative heatmaps, enhancing the model’s explainability. The proposed model fuses the local features of CNN model by explainable features of Grad-CAM and then feeds these enhanced local futures to a transformer in order to enrich global features. Finally, a Cascaded Upsampler (CUP) decoder for transforming encoded features to extract segmentation masks. Extensive experiments on the Kvasir and brain tumor MRI datasets demonstrate significant performance improvements over the state-of-the-art models across metrics such as the dice similarity coefficient (DSC), mean Intersection over Union (mIoU), precision, and recall. In The Kvasir Dataset, the DSC increased by 2.6%, while the IOU exhibited the most significant growth at 6%. Additionally, Recall saw an enhancement of 4.1%, and Precision improved by 3.4%. In the brain tumor MRI dataset, similar advancements were observed. The DSC rose by about 2%, and the IOU grew by 3.2%. Furthermore, Recall improved by 3%, and Precision showed a 1.1% increase.