MSGI3D: Multimodal semantic-geometric integration for 3D object detection in cluttered scenes
摘要
Sparse and incomplete point clouds frequently hinder robotic tasks such as navigation and object manipulation by obscuring object details and disrupting spatial relationships. While RGB images provide complementary information, existing methods struggle to fuse them effectively due to feature misalignment, inadequate small-scale object modeling, and insufficient global feature integration. To address these challenges, we propose MSGI3D, a novel framework that integrates 2D images with point cloud-based 3D object detection in sparse and occluded environments. The framework introduces a Synergistic Feature Fusion (SynerFusion) module to balance global semantic information from RGB images and geometric features of point clouds, ensuring robust multimodal fusion. To capture features across scales, the Context-driven Multi-scale Aggregation Framework (CMAF) improves detection of objects with diverse sizes and complexities. Additionally, the Linear Sparse Global Enhancement (LSGE) module employs sparse attention mechanisms to enhance global contextual representation in cluttered environments. Comprehensive experiments on SUN RGB-D, ScanNet V2, and S3DIS datasets validate the robustness of MSGI3D. Notably, it achieves a new state-of-the-art mAP@0.25 score of 70.07 on the SUN RGB-D dataset, outperforming existing methods and demonstrating its applicability across diverse scenarios.