3D Scene Reconstruction Using Lidar Point Clouds and Images
摘要
LiDAR technology has been a key component in the field of 3D mapping in recent years due to its relative advantages, such as high accuracy regardless of the environmental setting. Consequently, it has been utilized in numerous outdoor and indoor applications, including city planning, autonomous driving, and more . However, the adoption of LiDAR for 3D mapping is limited due to challenges such as the high cost of 3D LiDAR sensors, low resolution and refresh rate and speed shift of targets among others. To enhance and achieve a comprehensive 3D understanding of an environment, the use of multiple sensors in conjunction with LiDAR has been shown to yield significant results. The integration of LiDAR and cameras has demonstrated the ability to provide a rich context for 3D scene construction. Notable performance in various 3D vision tasks has been observed with these two common sensor types: LiDAR and cameras. Cameras provide scene context and semantic information, while LiDAR sensors deliver precise 3D geometry. However, due to the difference in data types—namely, image data and point clouds—a challenge exists in determining the best method to combine such data for the reconstruction task. A data processing pipeline using the convolutional occupancy networks framework has been developed, incorporating a PointNet encoder and a ResNet model for the image data to achieve scene reconstruction.