In the field of medical image segmentation, CNN-based models face challenges for learning distant contextual dependencies, while transformer-based approaches are limited by quadratic computational complexity. Mamba’s recent research demonstrates its ability to capture long-range interactions efficiently, with the advantage of linear computational complexity. Inspired by the Mamba architecture, U-Net skip connections, and the EMCAD decoder [1], We introduce an innovative architecture for skin lesion segmentation models, the Multi-Scale Vision Convolutional Attention U-Net (MVCA-UNet). To better model long-range dependencies, we embed the Vision State Space (VSS) module [2] within the encoder, enabling richer contextual feature extraction. In the decoder, we introduce the Multi-Scale Channel Convolution Block (MSCB), which uses depth-wise convolutions to capture complex spatial relationships and focus on salient regions. In addition, we propose a novel Multi-Scale Channel-Spatial Attention (MCSA) module in the skip connections, which simultaneously attends to both channel and spatial information. By combining detailed features from shallow layers with semantic cues from deeper layers, MCSA enhances segmentation accuracy. We conducted extensive experiments on the publicly available ISIC 2017 and ISIC 2018 skin lesion datasets, achieving accuracies of 96.81% and 95.65%, respectively. The results show that the MVCA-UNet model exhibits competitive performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MVCA-UNet: A Multi-scale Visual Convolutional Attention Architecture for Skin Lesion Segmentation

  • Zhaohui Wang,
  • Runzhi Xu,
  • Jiyong Xu,
  • Changfang Chen,
  • Ruixia Liu

摘要

In the field of medical image segmentation, CNN-based models face challenges for learning distant contextual dependencies, while transformer-based approaches are limited by quadratic computational complexity. Mamba’s recent research demonstrates its ability to capture long-range interactions efficiently, with the advantage of linear computational complexity. Inspired by the Mamba architecture, U-Net skip connections, and the EMCAD decoder [1], We introduce an innovative architecture for skin lesion segmentation models, the Multi-Scale Vision Convolutional Attention U-Net (MVCA-UNet). To better model long-range dependencies, we embed the Vision State Space (VSS) module [2] within the encoder, enabling richer contextual feature extraction. In the decoder, we introduce the Multi-Scale Channel Convolution Block (MSCB), which uses depth-wise convolutions to capture complex spatial relationships and focus on salient regions. In addition, we propose a novel Multi-Scale Channel-Spatial Attention (MCSA) module in the skip connections, which simultaneously attends to both channel and spatial information. By combining detailed features from shallow layers with semantic cues from deeper layers, MCSA enhances segmentation accuracy. We conducted extensive experiments on the publicly available ISIC 2017 and ISIC 2018 skin lesion datasets, achieving accuracies of 96.81% and 95.65%, respectively. The results show that the MVCA-UNet model exhibits competitive performance.