CCID: A Conditionally Controllable Image Diffusion Framework for Autonomous Driving Scene Generation
摘要
Generating realistic and diverse image data is crucial for enhancing perception and training reliability in autonomous driving. However, existing diffusion models struggle to effectively control viewpoint changes, resulting in poor generation controllability. This issue becomes more pronounced in unstructured road environments such as mining areas, where irregular terrain and sparse structures further amplify the problem. To address these challenges, we propose CCID (Conditionally Controllable Image Diffusion), a controllable generation framework that constructs prior frames based on point cloud conditions and utilizes new-view camera guidance to flexibly control the trajectory of generated data. The model conditions on the point cloud prior frame, uses the first frame from the original viewpoint as a reference image, and employs a pretrained UNet diffusion network for denoising and image synthesis. This design effectively incorporates geometric constraints into the diffusion process, producing images that are both high-quality and spatially consistent. Furthermore, an optimized prior-frame strategy is introduced to enhance controllability and stability in unstructured scenes with limited viewpoint coverage. Experiments on real-world driving data demonstrate that CCID outperforms existing methods in both controllability and image fidelity, achieving an FID of 51.6, far surpassing the 277.7 obtained by StreetGaussion.