RETRACTED CHAPTER: CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Ground Image Synthesis
摘要
Satellite-to-ground image synthesis aims at generating a realistic street-view image from its corresponding satellite-view image. Despite the considerable efforts towards improving the geometry structure of the synthetic street-view images, the texture fidelity and consistency have barely been explored and guaranteed in existing studies. In this work, we propose CrossViewDiff, a cross-view diffusion model for satellite-to-ground image synthesis. To address the challenges posed by the large discrepancy across views, we explore various types of conditional inputs and design a cross-view guided denoising process to enhance the texture consistency via pixel-wise texture alignment and blending as well as a masked cross-view attention scheme. Experimental results on two public cross-view benchmark datasets demonstrate that our CrossViewDiff generates high-quality street-view panoramas with more realistic structure and texture than current state-of-the-art for both rural and urban scenes. The code and models of this work will be made publicly available.