<p>Semantic segmentation is a fundamental task in computer vision. Over time, efforts have been made to mitigate the loss of spatial and global context information due to resolution reduction, starting with Fully Convolutional Networks (FCN). However, Convolutional Neural Networks (CNN) models still face challenges, as the down-sampling process reduces resolution, leading to the loss of essential spatial and global context information. In this paper, we combine the HRNet model, known for its exceptional performance in semantic segmentation, with the Channel and Spatial Attention (CSA) module. The CSA module highlights important spatial and global context information in semantic segmentation. It is structured so that the result of channel attention becomes the input for spatial attention. Channel attention recalibrates each channel’s resolution into a scalar value using a Squeeze-and-Excitation (SE) block, which is then multiplied by the original feature map. Spatial attention employs the Polarized Self-attention (PSA) module to compress the feature map channels from channel attention into a single resolution, and recalibrates them by multiplying by the input feature map. To assess the semantic segmentation performance of the proposed model, we compare and evaluate its performance with or without the CSA module, and against various models, using the LIP and PASCAL Context datasets. As a result of the experiment, in the LIP data set, mean Intersection over Union (mIoU) was improved by 0.5%, and Mean Pixel Accuracy (MPA) by 0.1%, improving both segmentation and class classification accuracy. In the PASCAL Context 59 data set, mIoU was maintained, while MPA was improved by 0.2%, which improved class classification accuracy. In the PASCAL Context 60 data set, mIoU decreased by 0.1%, while MPA was improved by 0.2%, improving class classification accuracy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CSA − HRNet: Channel and Spatial Attention High-Resolution Networks for Semantic Segmentation

  • Jin-Seong Kim,
  • Sung-Wook Park,
  • Hyun-Sung Yang,
  • Chun-Bo Sim,
  • Se-Hoon Jung

摘要

Semantic segmentation is a fundamental task in computer vision. Over time, efforts have been made to mitigate the loss of spatial and global context information due to resolution reduction, starting with Fully Convolutional Networks (FCN). However, Convolutional Neural Networks (CNN) models still face challenges, as the down-sampling process reduces resolution, leading to the loss of essential spatial and global context information. In this paper, we combine the HRNet model, known for its exceptional performance in semantic segmentation, with the Channel and Spatial Attention (CSA) module. The CSA module highlights important spatial and global context information in semantic segmentation. It is structured so that the result of channel attention becomes the input for spatial attention. Channel attention recalibrates each channel’s resolution into a scalar value using a Squeeze-and-Excitation (SE) block, which is then multiplied by the original feature map. Spatial attention employs the Polarized Self-attention (PSA) module to compress the feature map channels from channel attention into a single resolution, and recalibrates them by multiplying by the input feature map. To assess the semantic segmentation performance of the proposed model, we compare and evaluate its performance with or without the CSA module, and against various models, using the LIP and PASCAL Context datasets. As a result of the experiment, in the LIP data set, mean Intersection over Union (mIoU) was improved by 0.5%, and Mean Pixel Accuracy (MPA) by 0.1%, improving both segmentation and class classification accuracy. In the PASCAL Context 59 data set, mIoU was maintained, while MPA was improved by 0.2%, which improved class classification accuracy. In the PASCAL Context 60 data set, mIoU decreased by 0.1%, while MPA was improved by 0.2%, improving class classification accuracy.