Controlling Attention Map Better for Text-Guided Image Editing Diffusion Models
摘要
Large visual-language models such as GLIDE, DALL-E 2, and Imagen show remarkable capabilities in image generation. However, applying these models to image editing presents challenges as slight changes to the prompt can lead to vastly different final images. The emergence of text-guided diffusion models is gradually establishing itself as a mainstream technique for image generation due to its realism and versatility. The latent space noise in diffusion models possesses the ability to both preserve and modify the final image. Consequently, an increasing number of methods are employing diffusion models for text-guided image editing by exploring the process of noise generation. Despite the significant progress made by diffusion-based methods in image generation and editing, they often focus on specific application aspects. While these approaches yield outstanding results in their respective domains, there exists a lack of a unified method that integrates different editing approaches. This absence impedes users from choosing an algorithm that best suits their needs. To address this, we propose a convergent attention map modification framework. This framework seamlessly integrates various attention map control methods, allowing for controllable generation through different parameter combinations.