Accurate spatial perception is indispensable for robotic locomotion, with depth estimation playing a crucial role in enhancing spatial awareness. Recent advances have seen the integration of multimodal data becoming prevalent in robotic vision tasks, significantly improving precision and robustness in spatial perception. Despite the demonstrated potential of improving depth estimation in indoor settings through the fusion of visual and echoes data, aligning the transformed features from these modalities effectively continues to pose a challenge. In this paper, we propose a novel multimodal fusion approach that integrates visual and echo representations using a combination of self-attention and cross-attention mechanisms, enabling better alignment between features from both modalities. To further enhance the fusion of global semantic information, we introduce an innovative global multimodal fusion framework, which incorporates 3D scene context into the feature extraction layers across modalities, leading to improved depth estimation results. Comprehensive experiments conducted on the Replica and Matterport3D datasets, along with comparisons to state-of-the-art methods, demonstrate the effectiveness of the proposed approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-modal Scene Global Fusion Framework for Enhanced Depth Estimation

  • Anjie Wang,
  • Xujun Wei,
  • Mingxuan Chen,
  • Xiaoyan Jiang,
  • Yongbin Gao,
  • Zhijun Fang,
  • Siwei Ma

摘要

Accurate spatial perception is indispensable for robotic locomotion, with depth estimation playing a crucial role in enhancing spatial awareness. Recent advances have seen the integration of multimodal data becoming prevalent in robotic vision tasks, significantly improving precision and robustness in spatial perception. Despite the demonstrated potential of improving depth estimation in indoor settings through the fusion of visual and echoes data, aligning the transformed features from these modalities effectively continues to pose a challenge. In this paper, we propose a novel multimodal fusion approach that integrates visual and echo representations using a combination of self-attention and cross-attention mechanisms, enabling better alignment between features from both modalities. To further enhance the fusion of global semantic information, we introduce an innovative global multimodal fusion framework, which incorporates 3D scene context into the feature extraction layers across modalities, leading to improved depth estimation results. Comprehensive experiments conducted on the Replica and Matterport3D datasets, along with comparisons to state-of-the-art methods, demonstrate the effectiveness of the proposed approach.