<p>Real-time image semantic segmentation is a key technology for scene understanding and an important research direction in computer vision. It is essential for applications like autonomous driving, where quick response is critical. However, improving accuracy often sacrifices computational efficiency and inference speed. To address the trade-off among speed, model size, and accuracy in existing lightweight semantic segmentation models, we propose a multi-branch lightweight real-time semantic segmentation network. The network is designed based on MobileNetV2 and adopts a dual-branch structure to achieve efficient feature extraction. To enhance the context modeling ability of the context branch, a novel Dynamic Multi-Scale Context Fusion (DMSCF) module is proposed, which realizes adaptive multi-scale feature fusion by mathematically and dynamically adjusting the dilation rate. To improve the capture ability of spatial detail features, a Deformable Multi-modal Collaborative Attention (DMCA) is introduced using the deformable attention field and cross-modal collaborative learning. Furthermore, to effectively integrate the features from different branches, we designed a Feature Fusion Module (FFM) to ensure the rational and comprehensive integration of feature information. Although the dual-branch structure demonstrates excellent performance in segmentation tasks, the critical role of boundary information is often neglected. To address this issue, we propose the boundary auxiliary information prompt module, which is specifically designed to enhance the performance of detail segmentation. Experimental results indicate that our algorithm achieves segmentation accuracies of 78.3% and 77.6% on the Cityscapes and CamVid datasets, respectively, with inference speeds of 107 FPS and 114 FPS, respectively. These results successfully demonstrate an advanced balance between accuracy and real-time performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-mechanism collaborative segmentation network with boundary-aware prompt and dynamic attention enhancement

  • Jiaqi Zhong,
  • Huaming Qian

摘要

Real-time image semantic segmentation is a key technology for scene understanding and an important research direction in computer vision. It is essential for applications like autonomous driving, where quick response is critical. However, improving accuracy often sacrifices computational efficiency and inference speed. To address the trade-off among speed, model size, and accuracy in existing lightweight semantic segmentation models, we propose a multi-branch lightweight real-time semantic segmentation network. The network is designed based on MobileNetV2 and adopts a dual-branch structure to achieve efficient feature extraction. To enhance the context modeling ability of the context branch, a novel Dynamic Multi-Scale Context Fusion (DMSCF) module is proposed, which realizes adaptive multi-scale feature fusion by mathematically and dynamically adjusting the dilation rate. To improve the capture ability of spatial detail features, a Deformable Multi-modal Collaborative Attention (DMCA) is introduced using the deformable attention field and cross-modal collaborative learning. Furthermore, to effectively integrate the features from different branches, we designed a Feature Fusion Module (FFM) to ensure the rational and comprehensive integration of feature information. Although the dual-branch structure demonstrates excellent performance in segmentation tasks, the critical role of boundary information is often neglected. To address this issue, we propose the boundary auxiliary information prompt module, which is specifically designed to enhance the performance of detail segmentation. Experimental results indicate that our algorithm achieves segmentation accuracies of 78.3% and 77.6% on the Cityscapes and CamVid datasets, respectively, with inference speeds of 107 FPS and 114 FPS, respectively. These results successfully demonstrate an advanced balance between accuracy and real-time performance.