Multimodal Face Generation Model Based on Generative Adversarial Networks and Attention Mechanism
摘要
In this paper, we introduce a novel multimodal face generation model that seamlessly integrates generative adversarial networks (GANs) with a hybrid attention mechanism. The proposed architecture is designed to harness the power of GANs in generating realistic facial images while leveraging the attention mechanism to focus on salient features across different modalities. By incorporating a hybrid attention mechanism that combines cross-attention and self-attention, the proposed model enables us to effectively capture facial features from multiple modalities, resulting in more accurate and realistic generations. Our model’s effectiveness is demonstrated through its utilization of the CASIA-SURF, CASIA-SURF-CeFA and CelebA datasets, where the extensive results also reveal that the proposed model can attain top-tier accuracy while generating highly realistic outcomes. That indicates a significant step forward in the field of multimodal face generation and holds promise for various applications, such as face recognition, face manipulation, and facial animation.