Joint feature modulation mechanism for driving scene image synthesis by instance texture edge and spatial depth priors
摘要
In this paper, we study Conditional Image Synthesis (CIS) task towards producing photorealistic driving scenes, which plays a significant role in designing perception algorithms of autonomous driving vehicles. However, due to the sparsity of semantic condition representation and the lack of depth correlation, existing CIS methods face great challenges to produce high visual fidelity images with sharp details. To address this problem, this paper proposes a hybrid priors-assisted CIS solution based on the Joint Feature Modulation Mechanism (JFMM), which exploits complementary advantages of the instance texture edge and the spatial depth priors. Firstly, to synthesize rich instance details, JFMM adopts the edge-wise condition enhancement processing by channel-wise concatenating the original semantic layout and an aligned edge map as input of generator’s semantic normalization. It effectively improves the synthesis performance of filling fine appearance content within semantic regions. Second, to construct depth-of-field structure, we design a depth-adaptive normalization module to achieve a fine-grained 3D visual guidance by calculating the depth-wise normalization parameters of feature activations. It adequately supplements the pixel level depth correlations and enhances the generated stereoscopic content. In addition, we propose the driving scene-specific image quality assessment by two measurements of depth and edge similarity to validate the effectiveness. Experiments with quantitative and qualitative comparisons to state-of-the-art approaches demonstrate that the proposed method can achieve a competitive result on complex scene datasets Cityscapes and ADE20K-outdoor.