SM-GAN: Single-Stage and Multi-object Text Guided Image Editing
摘要
In recent years, text-guided scene image manipulation has received extensive attention in the computer vision community. Most of the existing research has focused on manipulating a single object from available conditioning information in each step. Manipulating multiple objects often relies on iterative networks progressively generating the edited image for each instruction. In this paper, we study a setting that allows users to edit multiple objects based on complex text instructions in a single stage and propose the Single-stage and Multi-object editing Generative Adversarial Network (SM-GAN) to tackle problems in this setting, which contains two key components: (i) the Spatial Semantic Enhancement module (SSE) deepens the spatial semantic prediction process to select all correct positions in the image space that need to be modified according to text instructions in a single stage, (ii) the Multi-object Detail Consistency module (MDC) learns semantic attributes adaptive modulation parameter conditioned on text instructions to effectively fuse text features and image features, and can ensure the generated visual attributes are aligned with text instructions. We construct the Multi-CLEVR dataset for CLEVR scene image construction using complex text instructions with single-stage processing. Extensive experiments on the Multi-CLEVR and CoDraw datasets have demonstrated the superior performance of the proposed method.