3DOPT-SLAM: Robust SLAM via Joint Optimization with 3D Object Pose Tracking
摘要
We propose 3DOPT-SLAM, a robust monocular SLAM method that leverages 3D object pose tracking to address feature scarcity in challenging environments. By incorporating known 3D object models as priors, any detectable object—regardless of its texture richness—can serve as a 6DoF pose anchor. Capitalizing on advances in object pose tracking, these methods robustly estimate object poses in real-time, even in textureless or weakly textured scenes. Based on these methods, we present a novel framework where point-based SLAM and 3D object poses mutually complement each other. Initially, camera localization is achieved solely through object tracking. Following the extraction of a sufficient number of feature point pairs, we establish a joint optimization framework between objects and feature points in both front-end and back-end threads. Simultaneously, we indicate that scene feature points can integrate with 3D object tracking constraints across multiple frames. The experiment results on the public dataset YCB-Video and TUM RGB-D cabinet sequence demonstrate significant improvements in the accuracy of camera trajectory estimation and the accuracy of object tracking. These results underscore that the collaboration between feature points and objects enhances camera localization and object tracking.