3DSSG-Cap: A Caption Enhanced Dataset for 3D Visual Grounding
摘要
With the rapid advancement of deep learning technologies and the availability of large-scale 3D point cloud datasets, 3D visual grounding tasks have garnered increasing attention in recent years. Although many contemporary studies have reported promising results, a notable challenge persists: most existing 3D visual grounding datasets rely heavily on human-written descriptions, which can be difficult to modify or extend. Furthermore, the quality of these descriptions often varies, impacting the consistency of models. In response to these issues, we introduce the 3DSSG-Cap dataset, which consists of 383,438 descriptions for 27,000 objects from 1,318 indoor scenes. Unlike traditional datasets, the descriptions in 3DSSG-Cap are generated using predefined templates, making them more flexible and easier to extend. In addition, we propose a novel method, 3DETRefer, to localize objects within the dataset. By integrating a transformer-based detector and a visual grounding fusion module, our approach significantly improves object localization and identification accuracy in complex 3D environments.