Vehiclesim: realistic and 3D-aware video editing with one image for autonomous driving
摘要
With the rapid development of autonomous driving technologies, the integration of controllable objects into videos for real-world simulation has drawn increasing attention. While state-of-the-art techniques like Neural Radiance Fields and 3D Gaussian Splatting are capable of rendering high-quality videos, they are often constrained by limited generalization abilities and substantial input source requirements, including multi-view images and LiDAR point clouds. Although 2D generation methods have been extensively explored for 3D generations, maintaining consistency remains a persistent challenge. To address these limitations, this paper proposes a two-stage video editing pipeline that generates videos featuring appearance- and 3D-controllable vehicles with fewer input sources. Our work introduces a mask estimation model in the first stage, which utilizes camera parameters to precisely manage the size and positioning of synthesized objects within videos. For the modification of the video editing model in the second stage, two innovative modules are presented. The reference encoder component of image-to-video diffusion models is enhanced by integrating a Selective Expert Module, significantly improving the fidelity of the generated objects to their reference images. We also creatively build a Multi-directional Mamba Module, which expands the capabilities of Mamba across multiple directions to ensure both spatial and temporal consistency in the generated videos. Extensive experiments prove that the presented method outperforms several leading models in the integration of realistic 3D-aware vehicles into images and videos. The manipulation capabilities of our solution facilitate the seamless generation of custom vehicles for driving simulations, effectively addressing the critical requirements for autonomous vehicle training and testing.