Multi-modal Scene Global Fusion Framework for Enhanced Depth Estimation
摘要
Accurate spatial perception is indispensable for robotic locomotion, with depth estimation playing a crucial role in enhancing spatial awareness. Recent advances have seen the integration of multimodal data becoming prevalent in robotic vision tasks, significantly improving precision and robustness in spatial perception. Despite the demonstrated potential of improving depth estimation in indoor settings through the fusion of visual and echoes data, aligning the transformed features from these modalities effectively continues to pose a challenge. In this paper, we propose a novel multimodal fusion approach that integrates visual and echo representations using a combination of self-attention and cross-attention mechanisms, enabling better alignment between features from both modalities. To further enhance the fusion of global semantic information, we introduce an innovative global multimodal fusion framework, which incorporates 3D scene context into the feature extraction layers across modalities, leading to improved depth estimation results. Comprehensive experiments conducted on the Replica and Matterport3D datasets, along with comparisons to state-of-the-art methods, demonstrate the effectiveness of the proposed approach.