In cardiac image segmentation tasks, current model architectures face multiple challenges, including insufficient local feature capture, inadequate multi-scale information representation, and artifacts caused by standard upsampling techniques. To address these challenges, we propose the ​Residual-Swin Multi-Scale Attention Network (RSMA-Net), a novel method that innovatively integrates residual fusion of ResNet and Swin Transformer with multi-scale reconstruction. In the encoder, RSMA-Net leverages the multi-scale feature extraction capabilities of ResNet and Swin Transformer. By performing multi-scale residual fusion on outputs from different ResNet layers and subsequently conducting global modeling of the fused features at each layer, the architecture effectively combines CNN’s local feature capture ability with Transformer’s dependency modeling strengths. To expand the receptive field and capture richer contextual features in the decoder, we design the ​Dual Multi-Scale Depthwise Convolution Block (DMDCB) for effective learning and reconstruction of complex cardiac structural features. Additionally, we introduce the ​Local-Global Patch Attention Module (LGPAM) to fuse local and global features, employing channel and spatial attention mechanisms to emphasize critical feature representations. To mitigate artifacts from traditional upsampling techniques, we propose the ​Residual Up-Sampling method (ResUpSam), which preserves more low-level features through transposed convolution combined with residual connections, ensuring the integrity of fine structures. Experimental results show RSMA-Net achieves promising performance on cardiac imaging datasets, especially in accurate boundary delineation and detailed internal structure segmentation. This work provides an efficient and innovative solution for cardiac image segmentation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

RSMA-Net:ResNet-Swin Residual Cooperative Encoding and Multi-scale Attentional Reconstruction Network for Cardiac Boundary Segmentation

  • Yixiang Du,
  • Shuaishuai Zhang,
  • Ruixia Liu,
  • Pengyao Xu

摘要

In cardiac image segmentation tasks, current model architectures face multiple challenges, including insufficient local feature capture, inadequate multi-scale information representation, and artifacts caused by standard upsampling techniques. To address these challenges, we propose the ​Residual-Swin Multi-Scale Attention Network (RSMA-Net), a novel method that innovatively integrates residual fusion of ResNet and Swin Transformer with multi-scale reconstruction. In the encoder, RSMA-Net leverages the multi-scale feature extraction capabilities of ResNet and Swin Transformer. By performing multi-scale residual fusion on outputs from different ResNet layers and subsequently conducting global modeling of the fused features at each layer, the architecture effectively combines CNN’s local feature capture ability with Transformer’s dependency modeling strengths. To expand the receptive field and capture richer contextual features in the decoder, we design the ​Dual Multi-Scale Depthwise Convolution Block (DMDCB) for effective learning and reconstruction of complex cardiac structural features. Additionally, we introduce the ​Local-Global Patch Attention Module (LGPAM) to fuse local and global features, employing channel and spatial attention mechanisms to emphasize critical feature representations. To mitigate artifacts from traditional upsampling techniques, we propose the ​Residual Up-Sampling method (ResUpSam), which preserves more low-level features through transposed convolution combined with residual connections, ensuring the integrity of fine structures. Experimental results show RSMA-Net achieves promising performance on cardiac imaging datasets, especially in accurate boundary delineation and detailed internal structure segmentation. This work provides an efficient and innovative solution for cardiac image segmentation.