<p>We present Coordinate-Attention Shape-Attentive U-Net (CA-SAUNet), a lightweight segmentation architecture that replaces the Squeeze-and-Excitation (SE) blocks in Shape-Attentive UNet (SAUNet) with position-aware Coordinate Attention (CA) modules. By embedding directional context into channel re-weighting, CA-SAUNet simultaneously learns what to attend to and where it appears, while retaining SAUNet’s gated shape stream for boundary refinement. Ablation experiments confirm that this single design change, requiring no extra parameters, yields consistently sharper masks than SE- or Convolutional Block Attention Module (CBAM) based decoders. On the International Skin Imaging Collaboration (ISIC) 2018 skin-lesion benchmark, CA-SAUNet achieves 0.847 mean Intersection-over-Union (mIoU) and 0.911 Dice, surpassing the previous state of the art with only 33&#xa0;M parameters. Cross-domain tests on our pixel-annotated PlantDoc set further demonstrate robustness: CA-SAUNet attains 0.863 mIoU and 0.912 Dice, outperforming vanilla UNet and attention-based baselines despite cluttered field imagery. To ensure reliability, these results were validated using a rigorous 5-fold cross-validation protocol with strict separation of training and testing samples. The network maintains a compact 136&#xa0;MB weight footprint, which can be further reduced to 68&#xa0;MB via half-precision quantization, making it highly suitable for real-time deployment on clinical workstations or resource-constrained edge devices. Comprehensive comparisons confirm that the proposed coordinate attention delivers the highest overlap scores on both datasets without sacrificing inference speed. The achieved results enhance diagnostic accuracy in medical imaging and improve disease detection in agriculture, demonstrating the practical benefits of position-sensitive attention.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CA-SAUNet: Coordinate-Attention Enhanced Shape Attentive UNet for medical and agricultural image segmentation

  • Gültekin Işık,
  • Fesih Keskin

摘要

We present Coordinate-Attention Shape-Attentive U-Net (CA-SAUNet), a lightweight segmentation architecture that replaces the Squeeze-and-Excitation (SE) blocks in Shape-Attentive UNet (SAUNet) with position-aware Coordinate Attention (CA) modules. By embedding directional context into channel re-weighting, CA-SAUNet simultaneously learns what to attend to and where it appears, while retaining SAUNet’s gated shape stream for boundary refinement. Ablation experiments confirm that this single design change, requiring no extra parameters, yields consistently sharper masks than SE- or Convolutional Block Attention Module (CBAM) based decoders. On the International Skin Imaging Collaboration (ISIC) 2018 skin-lesion benchmark, CA-SAUNet achieves 0.847 mean Intersection-over-Union (mIoU) and 0.911 Dice, surpassing the previous state of the art with only 33 M parameters. Cross-domain tests on our pixel-annotated PlantDoc set further demonstrate robustness: CA-SAUNet attains 0.863 mIoU and 0.912 Dice, outperforming vanilla UNet and attention-based baselines despite cluttered field imagery. To ensure reliability, these results were validated using a rigorous 5-fold cross-validation protocol with strict separation of training and testing samples. The network maintains a compact 136 MB weight footprint, which can be further reduced to 68 MB via half-precision quantization, making it highly suitable for real-time deployment on clinical workstations or resource-constrained edge devices. Comprehensive comparisons confirm that the proposed coordinate attention delivers the highest overlap scores on both datasets without sacrificing inference speed. The achieved results enhance diagnostic accuracy in medical imaging and improve disease detection in agriculture, demonstrating the practical benefits of position-sensitive attention.