<p>Extensive research has focused on developing efficient and accurate solutions for the critical task of medical image segmentation. Approaches have evolved from hand-crafted pipelines to deep convolutional neural networks (CNNs), and more recently, to Transformer-based hybrid models. Among these, hierarchical encoder–decoder architectures remain prevalent, where skip connections are crucial in transmitting spatial features from encoders to decoders. However, conventional skip connections operate in static and passive modes, and cannot adaptively fuse multi-scale features or capture semantic relationships across resolution levels. Although attention-based skip enhancements have been proposed, they are often architecture-specific and difficult to generalize. In this study, we propose TransSkip, a novel transformer-based skip connection module that embeds both self-attention and cross-attention directly within the skip path. This enables dynamic and learnable multi-scale feature fusion across encoder levels, transforming skip connections into active semantic reasoning pathways. TransSkip is modular and architecture agnostic, supporting seamless integration with a range of hierarchical encoder–decoder networks, including CNN-based, Transformer-based, and hybrid models. Extensive experiments across 2D and 3D datasets (BUSI, Kvasir-SEG, MSD-Spleen) and multiple network backbones (U-Net, TransUNet, TransAttUNet, MCV-UNet) demonstrate that TransSkip consistently improves segmentation accuracy, with statistically significant gains and minimal parameter overhead. These results highlight the potential of TransSkip as a generalizable and efficient architectural enhancement for medical image segmentation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Harnessing transformer-based attention mechanisms for multi-scale feature fusion in medical image segmentation

  • Rabeea Fatma Khan,
  • Mu Sook Lee,
  • Byoung-Dai Lee

摘要

Extensive research has focused on developing efficient and accurate solutions for the critical task of medical image segmentation. Approaches have evolved from hand-crafted pipelines to deep convolutional neural networks (CNNs), and more recently, to Transformer-based hybrid models. Among these, hierarchical encoder–decoder architectures remain prevalent, where skip connections are crucial in transmitting spatial features from encoders to decoders. However, conventional skip connections operate in static and passive modes, and cannot adaptively fuse multi-scale features or capture semantic relationships across resolution levels. Although attention-based skip enhancements have been proposed, they are often architecture-specific and difficult to generalize. In this study, we propose TransSkip, a novel transformer-based skip connection module that embeds both self-attention and cross-attention directly within the skip path. This enables dynamic and learnable multi-scale feature fusion across encoder levels, transforming skip connections into active semantic reasoning pathways. TransSkip is modular and architecture agnostic, supporting seamless integration with a range of hierarchical encoder–decoder networks, including CNN-based, Transformer-based, and hybrid models. Extensive experiments across 2D and 3D datasets (BUSI, Kvasir-SEG, MSD-Spleen) and multiple network backbones (U-Net, TransUNet, TransAttUNet, MCV-UNet) demonstrate that TransSkip consistently improves segmentation accuracy, with statistically significant gains and minimal parameter overhead. These results highlight the potential of TransSkip as a generalizable and efficient architectural enhancement for medical image segmentation.