<p>Introducing a visual Simultaneous Localization and Mapping (SLAM) algorithm designed for mobile robots in environments containing numerous moving objects. It addresses the challenges these dynamic scenes pose to the accuracy and reliability of long-term localization. The algorithm embeds a lightweight target detection network in the front-end of the ORB-SLAM3 system for detecting dynamic targets, replaces its backbone network with the MobileNetV3 model, and enhances feature extraction capability by integrating improved Spatial Pyramid Pooling Fusion and Context Module (SPPFCM) and Coordinate Attention (CA) mechanisms. Experimental results on VOC2007+VOC2012 datasets demonstrate that the improved model reduces the number of network parameters by 32.2<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11370_2025_625_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation> compared to the original YOLOv5s model while maintaining equivalent accuracy. The fused visual odometry with target detection is combined with the Lucas-Kanade (LK) optical flow method in the tracking thread to achieve dynamic feature point rejection, retaining only static feature points for subsequent feature matching and camera pose estimation. To mitigate the impact of insufficient feature matches after dynamic feature removal, a more accurate Box Average Difference (BAD) replaces the original Binary Robust Independent Elementary Features (BRIEF) descriptor. Experiments on the TUM dataset show that compared with ORB-SLAM3, the proposed algorithm demonstrates a 14.59<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11370_2025_625_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation> improvement in feature matching accuracy and achieves a 95.1<InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11370_2025_625_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation> enhancement in camera positioning accuracy under high-dynamic environments, thereby significantly enhancing the system’s positioning precision and robustness in dynamic scenarios.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Visual SLAM and dense map reconstruction in highly dynamic environments

  • Zengzhen Mi,
  • Xingyuan Xu,
  • Jing Tang,
  • Chengzhi Leiyu,
  • Yifu Chen

摘要

Introducing a visual Simultaneous Localization and Mapping (SLAM) algorithm designed for mobile robots in environments containing numerous moving objects. It addresses the challenges these dynamic scenes pose to the accuracy and reliability of long-term localization. The algorithm embeds a lightweight target detection network in the front-end of the ORB-SLAM3 system for detecting dynamic targets, replaces its backbone network with the MobileNetV3 model, and enhances feature extraction capability by integrating improved Spatial Pyramid Pooling Fusion and Context Module (SPPFCM) and Coordinate Attention (CA) mechanisms. Experimental results on VOC2007+VOC2012 datasets demonstrate that the improved model reduces the number of network parameters by 32.2 \(\%\) % compared to the original YOLOv5s model while maintaining equivalent accuracy. The fused visual odometry with target detection is combined with the Lucas-Kanade (LK) optical flow method in the tracking thread to achieve dynamic feature point rejection, retaining only static feature points for subsequent feature matching and camera pose estimation. To mitigate the impact of insufficient feature matches after dynamic feature removal, a more accurate Box Average Difference (BAD) replaces the original Binary Robust Independent Elementary Features (BRIEF) descriptor. Experiments on the TUM dataset show that compared with ORB-SLAM3, the proposed algorithm demonstrates a 14.59 \(\%\) % improvement in feature matching accuracy and achieves a 95.1 \(\%\) % enhancement in camera positioning accuracy under high-dynamic environments, thereby significantly enhancing the system’s positioning precision and robustness in dynamic scenarios.