Synthesizing driving views is crucial for extending training data in autonomous driving scenes. Recently, Neural Radiance Fields (NeRF) have achieved impressive results in novel view synthesis tasks for bounded scenes. However, due to implicit inconsistency, most existing NeRF-like models face performance degradation when reconstructing autonomous driving scenes. In this paper, we present StereoNeRF, which leverages the characteristics of stereo cameras to mitigate the negative effects of geometric uncertainty in volume rendering and enhance the performance of synthesized views. Firstly, we propose a novel loss term to regularize implicit geometric consistency by exploiting the photometric consistency between image pairs captured by stereo cameras. Furthermore, we introduce a data augmentation method to generate views across image pairs from stereo cameras based on the tracks of key points provided by Structure-from-Motion (SfM), which helps address performance degradation caused by shape-radiance ambiguity. Experiments on the KITTI-360 dataset demonstrate that our approach synthesizes photo-realistic novel views in autonomous driving scenes.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

StereoNeRF: Learning Radiance Fields from Stereo Observation for Driving View Synthesis

  • Dexin Qi,
  • Zhihong Zhang,
  • Tao Tao,
  • Xuesong Mei

摘要

Synthesizing driving views is crucial for extending training data in autonomous driving scenes. Recently, Neural Radiance Fields (NeRF) have achieved impressive results in novel view synthesis tasks for bounded scenes. However, due to implicit inconsistency, most existing NeRF-like models face performance degradation when reconstructing autonomous driving scenes. In this paper, we present StereoNeRF, which leverages the characteristics of stereo cameras to mitigate the negative effects of geometric uncertainty in volume rendering and enhance the performance of synthesized views. Firstly, we propose a novel loss term to regularize implicit geometric consistency by exploiting the photometric consistency between image pairs captured by stereo cameras. Furthermore, we introduce a data augmentation method to generate views across image pairs from stereo cameras based on the tracks of key points provided by Structure-from-Motion (SfM), which helps address performance degradation caused by shape-radiance ambiguity. Experiments on the KITTI-360 dataset demonstrate that our approach synthesizes photo-realistic novel views in autonomous driving scenes.