Place recognition stands as a pivotal research area within computer vision and artificial intelligence, offering substantial application potential. It serves as a foundational element in simultaneous localization and mapping systems and robot global positioning. While existing multimodal fusion techniques have achieved high recall rates in autonomous driving applications, their performance can suffer from changes in viewpoint or scene variability. This study introduces CVMPR, a novel approach specifically tailored for aerial data in the domain of drones. To enhance the network’s generalization and robustness, we integrate projection-based and point-based methods for capturing global features from environmental point clouds. Concurrently, a residual neural network extracts detailed local features from images. The key innovation lies in a cross-attention transformer designed to fuse information from different modalities effectively. This transformer not only integrates complementary data but also preserves original features, thereby maximizing correlations between images and point clouds. Our methodology undergoes rigorous evaluation on two established benchmark datasets, Oxford RobotCar and MUN-FRL. Experimental results unequivocally demonstrate that CVMPR surpasses state-of-the-art methods in terms of generalization and robustness, significantly improving the accuracy of cross-view place recognition.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-View Multimodal Place Recognition

  • Lu Xu,
  • Shuaixin Li,
  • Xiaozhou Zhu,
  • Wen Yao

摘要

Place recognition stands as a pivotal research area within computer vision and artificial intelligence, offering substantial application potential. It serves as a foundational element in simultaneous localization and mapping systems and robot global positioning. While existing multimodal fusion techniques have achieved high recall rates in autonomous driving applications, their performance can suffer from changes in viewpoint or scene variability. This study introduces CVMPR, a novel approach specifically tailored for aerial data in the domain of drones. To enhance the network’s generalization and robustness, we integrate projection-based and point-based methods for capturing global features from environmental point clouds. Concurrently, a residual neural network extracts detailed local features from images. The key innovation lies in a cross-attention transformer designed to fuse information from different modalities effectively. This transformer not only integrates complementary data but also preserves original features, thereby maximizing correlations between images and point clouds. Our methodology undergoes rigorous evaluation on two established benchmark datasets, Oxford RobotCar and MUN-FRL. Experimental results unequivocally demonstrate that CVMPR surpasses state-of-the-art methods in terms of generalization and robustness, significantly improving the accuracy of cross-view place recognition.