Image features typically contain rich structured semantics, with sub-regions interconnected through intricate semantic relationships. Similarly, words in text descriptions are linked through semantic connections. However, as networks deepen, these image semantics often weaken, leading to a loss of coherence in image synthesis. Moreover, existing methods often overlook the structured semantics between words, making it challenging to align image sub-regions with isolated word embeddings. To address these challenges, we propose an Efficient Attention-Bridged Fusion GAN (EAFGAN) for high-quality image synthesis. EAFGAN introduces two novel components: a Multi-Scale Dilated Fusion Module (MDFM), which enhances the capture and representation of complex semantic structures in images, and a Word-Image Cross Fusion Module (WCFM), which enriches word embeddings with multi-scale features to guide the structured generation of images. Extensive experimental results and ablation studies demonstrate that our proposed EAFGAN can effectively improve the quality of the generated images.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient Attention-Bridged Fusion GAN for Text-to-Image Synthesis

  • Junpeng Liu,
  • Qingfeng Wu

摘要

Image features typically contain rich structured semantics, with sub-regions interconnected through intricate semantic relationships. Similarly, words in text descriptions are linked through semantic connections. However, as networks deepen, these image semantics often weaken, leading to a loss of coherence in image synthesis. Moreover, existing methods often overlook the structured semantics between words, making it challenging to align image sub-regions with isolated word embeddings. To address these challenges, we propose an Efficient Attention-Bridged Fusion GAN (EAFGAN) for high-quality image synthesis. EAFGAN introduces two novel components: a Multi-Scale Dilated Fusion Module (MDFM), which enhances the capture and representation of complex semantic structures in images, and a Word-Image Cross Fusion Module (WCFM), which enriches word embeddings with multi-scale features to guide the structured generation of images. Extensive experimental results and ablation studies demonstrate that our proposed EAFGAN can effectively improve the quality of the generated images.