<p>Currently, most mainstream visual SLAM (Simultaneous Localization and Mapping) algorithms are based on the assumption of a static environment, and this strong assumption limits the application of the system in highly dynamic and realistic environments. With the rapid development of deep learning, semantic SLAM combined with deep learning has emerged as a primary solution to handle dynamic scenes. However, this approach faces challenges related to computational resource demands and model complexity. To address these limitations, this paper proposes a lightweight semantic segmentation network, LSSMask (<i>Lightweight Semantic Segmentation Network</i>), which not only reduces the parameters of the semantic segmentation model but also enhances the representational capacity of features. Firstly, we design a lightweight GDS-ECA convolution (<i>Ghost-Depthwise Separable Convolution with Efficient Channel Attention Block</i>), which replaces traditional convolution with depthwise separable convolution to reduce both parameters and computational costs, while incorporating the ECA (<i>Efficient Channel Attention</i>) mechanism to strengthen the feature representation. Second, the BGTNet feature extraction network (<i>Bottleneck GDS-ECA Convolution with Transformer Network</i>) is introduced, applying GDS-ECA convolution to the neck module’s convolutional layers, replacing standard convolution operations. Additionally, we replace the traditional convolutions used for extracting high-level semantics in the FPN (<i>Feature Pyramid Network</i>) with the proposed GDS-ECA convolution, constructing the Lightweight Feature Pyramid Network (<i>L-FPN</i>). Together, BGTNet and L-FPN form the backbone of the LSSMask network. Finally, experimental validation on the COCO dataset demonstrates that the proposed model effectively improves both the efficiency and performance of the network, while maintaining segmentation accuracy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LSSMask: a lightweight semantic segmentation network for dynamic object

  • Xiaofeng Lian,
  • Maomao Kang,
  • Li Tan,
  • Xiao Sun,
  • Yanli Wang

摘要

Currently, most mainstream visual SLAM (Simultaneous Localization and Mapping) algorithms are based on the assumption of a static environment, and this strong assumption limits the application of the system in highly dynamic and realistic environments. With the rapid development of deep learning, semantic SLAM combined with deep learning has emerged as a primary solution to handle dynamic scenes. However, this approach faces challenges related to computational resource demands and model complexity. To address these limitations, this paper proposes a lightweight semantic segmentation network, LSSMask (Lightweight Semantic Segmentation Network), which not only reduces the parameters of the semantic segmentation model but also enhances the representational capacity of features. Firstly, we design a lightweight GDS-ECA convolution (Ghost-Depthwise Separable Convolution with Efficient Channel Attention Block), which replaces traditional convolution with depthwise separable convolution to reduce both parameters and computational costs, while incorporating the ECA (Efficient Channel Attention) mechanism to strengthen the feature representation. Second, the BGTNet feature extraction network (Bottleneck GDS-ECA Convolution with Transformer Network) is introduced, applying GDS-ECA convolution to the neck module’s convolutional layers, replacing standard convolution operations. Additionally, we replace the traditional convolutions used for extracting high-level semantics in the FPN (Feature Pyramid Network) with the proposed GDS-ECA convolution, constructing the Lightweight Feature Pyramid Network (L-FPN). Together, BGTNet and L-FPN form the backbone of the LSSMask network. Finally, experimental validation on the COCO dataset demonstrates that the proposed model effectively improves both the efficiency and performance of the network, while maintaining segmentation accuracy.