<p>In the field of semantic segmentation, the limited receptive field of convolutional neural networks leads to insufficient extraction of global features, thereby affecting the accuracy of network segmentation. To address this issue, a Hierarchical Hybrid Encoder Network (HHEnet) based on Transformers is proposed for semantic segmentation. Firstly, to solve the problem of limited global feature information caused by the network’s limited receptive field, a Hierarchical Hybrid Encoder (HHE) is introduced, which consists of a Hierarchical Convolutional Encoder (HCE) and a Hierarchical Transformer Encoder (HTE). The encoder combines the advantages of convolution and transformers, allowing for effective extraction of both shallow and deep features. In order to further enhance spatial and global semantic information, the Feature Enhancement Module (FEM) was introduced, which consisted of two feature enhancement modules: spatial feature enhancement module (SEM) and global feature enhancement module (GEM), which enhanced spatial detail information and global semantic information respectively. Thus the accuracy of semantic segmentation can be improved. Finally, to alleviate the discrepancy between the features of the convolutional encoder and the transformer encoder, a Feature Guidance Module (FGM) is introduced. Experimental results conducted on Cityscapes, ADE20K and PASCAL VOC2012 datasets achieved mIoU scores of up to 81.9%, 49.4% and 79.1%, respectively. Compared to state-of-the-art networks, the research results confirm the higher segmentation accuracy of the proposed HHEnet in this study.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Transformer-Based Hierarchical Hybrid Encoder Network for Semantic Segmentation

  • Shan Zhao,
  • Xuan Wu,
  • Kaiwen Tian,
  • Yang Yuan

摘要

In the field of semantic segmentation, the limited receptive field of convolutional neural networks leads to insufficient extraction of global features, thereby affecting the accuracy of network segmentation. To address this issue, a Hierarchical Hybrid Encoder Network (HHEnet) based on Transformers is proposed for semantic segmentation. Firstly, to solve the problem of limited global feature information caused by the network’s limited receptive field, a Hierarchical Hybrid Encoder (HHE) is introduced, which consists of a Hierarchical Convolutional Encoder (HCE) and a Hierarchical Transformer Encoder (HTE). The encoder combines the advantages of convolution and transformers, allowing for effective extraction of both shallow and deep features. In order to further enhance spatial and global semantic information, the Feature Enhancement Module (FEM) was introduced, which consisted of two feature enhancement modules: spatial feature enhancement module (SEM) and global feature enhancement module (GEM), which enhanced spatial detail information and global semantic information respectively. Thus the accuracy of semantic segmentation can be improved. Finally, to alleviate the discrepancy between the features of the convolutional encoder and the transformer encoder, a Feature Guidance Module (FGM) is introduced. Experimental results conducted on Cityscapes, ADE20K and PASCAL VOC2012 datasets achieved mIoU scores of up to 81.9%, 49.4% and 79.1%, respectively. Compared to state-of-the-art networks, the research results confirm the higher segmentation accuracy of the proposed HHEnet in this study.