Text-to-image generation combines NLP and computer vision but faces challenges. (1) Invalid feature interference: Models may learn ineffective features from training data, affecting generated images. (2) Semantic alignment: Capturing and expressing visual details such as color, shape, and spatial relationships is difficult due to cross-modal mapping. This paper proposes GCAttn-GAN, a text-to-image model leveraging gated convolutional attention. The model extracts fused features and dynamically adjusts their importance, enhancing feature representation in convolutional neural networks. Its convolutional attention module integrates text semantics with image feature channels, extracting high-quality features to improve training and mitigate invalid feature interference (1). To enhance text-image alignment, a recurrent consistency network generates text from the final image, minimizing cross-entropy loss with the original text. This improves semantic consistency between text and image content, addressing (2). The model follows a three-stage image generation process: the first two stages establish structure and color distribution, while the third refines quality. Experiments on CUB-200-2011 and COCO datasets, including quantitative, visual, and ablation analyses, demonstrate the model’s effectiveness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Text Generation Image Model Based on Gated Convolution Attention Generation Adversarial Network

  • Meiling Liu,
  • Wenlong Chen,
  • Longchang Liang,
  • Jiyun Zhou

摘要

Text-to-image generation combines NLP and computer vision but faces challenges. (1) Invalid feature interference: Models may learn ineffective features from training data, affecting generated images. (2) Semantic alignment: Capturing and expressing visual details such as color, shape, and spatial relationships is difficult due to cross-modal mapping. This paper proposes GCAttn-GAN, a text-to-image model leveraging gated convolutional attention. The model extracts fused features and dynamically adjusts their importance, enhancing feature representation in convolutional neural networks. Its convolutional attention module integrates text semantics with image feature channels, extracting high-quality features to improve training and mitigate invalid feature interference (1). To enhance text-image alignment, a recurrent consistency network generates text from the final image, minimizing cross-entropy loss with the original text. This improves semantic consistency between text and image content, addressing (2). The model follows a three-stage image generation process: the first two stages establish structure and color distribution, while the third refines quality. Experiments on CUB-200-2011 and COCO datasets, including quantitative, visual, and ablation analyses, demonstrate the model’s effectiveness.