Achieving robust localization performance in dynamic environments poses a significant challenge for current visual SLAM systems. Most existing research tackles this issue by employing either semantic segmentation-based methods or geometric approaches to filter out dynamic objects. However, the former often demands excessive computational resources, while the latter tends to falter in highly dynamic scenarios. To address these challenges, this paper proposes a method for dynamic environment recognition that integrates efficient object detection networks with optical flow estimation. Moreover, it identifies dynamic objects while retaining pseudo-dynamic objects that are actually stationary. Additionally, it utilizes the disparity between the depth values of dynamic objects and the background in depth images for clustering, yielding pseudo-semantic masks. This approach achieves efficient removal of dynamic objects akin to semantic segmentation. Experimental results on publicly available datasets demonstrate that the proposed method yields comparable or even superior performance to semantic segmentation-based techniques, with higher computational efficiency and minimal resource requirements. The average inference time was 53.4ms per frame, enabling real-time operation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Real-Time RGB-D Odometry in Dynamic Environment Based on Object Detection and Depth Clustering

  • Baifan Chen,
  • Jiadong Wan,
  • Yuxuan Jin,
  • Chengyu Peng

摘要

Achieving robust localization performance in dynamic environments poses a significant challenge for current visual SLAM systems. Most existing research tackles this issue by employing either semantic segmentation-based methods or geometric approaches to filter out dynamic objects. However, the former often demands excessive computational resources, while the latter tends to falter in highly dynamic scenarios. To address these challenges, this paper proposes a method for dynamic environment recognition that integrates efficient object detection networks with optical flow estimation. Moreover, it identifies dynamic objects while retaining pseudo-dynamic objects that are actually stationary. Additionally, it utilizes the disparity between the depth values of dynamic objects and the background in depth images for clustering, yielding pseudo-semantic masks. This approach achieves efficient removal of dynamic objects akin to semantic segmentation. Experimental results on publicly available datasets demonstrate that the proposed method yields comparable or even superior performance to semantic segmentation-based techniques, with higher computational efficiency and minimal resource requirements. The average inference time was 53.4ms per frame, enabling real-time operation.