<p>Medical image segmentation is a crucial technology for disease diagnosis and treatment planning. However, current approaches face challenges in capturing global semantic dependencies and integrating cross-layer features. While Convolutional Neural Networks (CNNs) excel at extracting local features, they struggle with long-range dependencies; Transformers effectively model global context but may compromise spatial details. To address these limitations, this paper proposes a novel hybrid CNN–Transformer architecture, Dual Attention and Cross-layer Fusion Network (DCF-Net). Based on an encoder–decoder framework, DCF-Net introduces two key modules: the Channel-Adaptive Sparse Attention (CASA) module and the Synergistic Skip-connection and Cross-layer Fusion (SSCF) module. Specifically, CASA enhances semantic modeling by filtering critical features and focusing on anatomically important regions, while SSCF enables effective hierarchical feature fusion by bridging encoder–decoder representations. Extensive experiments on the Synapse, ACDC, and ISIC2017 datasets demonstrate that DCF-Net achieves state-of-the-art performance without pre-training. This work highlights the value of cross-layer fusion and attention mechanism, providing a robust and generalizable solution for medical image segmentation tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A dual attention and cross layer fusion network with a hybrid CNN and transformer architecture for medical image segmentation

  • Jiahong Chen,
  • Zhengyou Liang,
  • Xiangyan Lu

摘要

Medical image segmentation is a crucial technology for disease diagnosis and treatment planning. However, current approaches face challenges in capturing global semantic dependencies and integrating cross-layer features. While Convolutional Neural Networks (CNNs) excel at extracting local features, they struggle with long-range dependencies; Transformers effectively model global context but may compromise spatial details. To address these limitations, this paper proposes a novel hybrid CNN–Transformer architecture, Dual Attention and Cross-layer Fusion Network (DCF-Net). Based on an encoder–decoder framework, DCF-Net introduces two key modules: the Channel-Adaptive Sparse Attention (CASA) module and the Synergistic Skip-connection and Cross-layer Fusion (SSCF) module. Specifically, CASA enhances semantic modeling by filtering critical features and focusing on anatomically important regions, while SSCF enables effective hierarchical feature fusion by bridging encoder–decoder representations. Extensive experiments on the Synapse, ACDC, and ISIC2017 datasets demonstrate that DCF-Net achieves state-of-the-art performance without pre-training. This work highlights the value of cross-layer fusion and attention mechanism, providing a robust and generalizable solution for medical image segmentation tasks.