AMM-GAN: Attribute-Matching Memory for Person Text-to-Image Generation
摘要
Using GANs to generate realistic images from text descriptions has achieved great progress. However, it is difficult to represent attributes such as gender, clothing, hairstyle, and accessories completely and accurately, so generating full-body person images remains a challenge. In this paper, we propose the Attribute-Matching Memory Generative Adversarial Network(AMM-GAN) to generate high-quality full-body person images directly from the text. The AMM-GAN uses an attribute text feature extraction module to extract relevant attribute information from the text, and the encoder is dynamically updated to extract more accurate features. We also propose an attribute-matching memory (AMM) generator to refine the image by improving the attention to the whole attribute in a memory network composed of image and text information. Furthermore, the real-result-driven discriminators guide the network to produce more realistic and natural images and achieve better image-text matching. Experimental results demonstrate that our AMM-GAN achieves the best performance in FID and R-Precison metrics compared to existing models.