Visual Street Localization Refinement Method Using Differentiable Rendering
摘要
Estimating the accurate location and pose of the camera on a vehicle is important for AR driving applications. In this work, we use a single image input and GIS map information to refine the raw geo-location and orientation from sensors on urban streets. We parse the urban scene in the image and obtain semantic and depth cues of buildings in the image using state-of-the-art deep learning methods. A 2.5D map is used as a reference to render the corresponding virtual semantic and depth maps given the initial sensor input. A camera pose optimization is performed by taking advantage of differentiable rendering to achieve maximal matching between the real image and the rendering. The results show that our method performs well on the Cityscapes dataset.