<p>Facial image synthesis is a pivotal research domain with wide-ranging applications in facial recognition, virtual reality, and the entertainment industry. This paper introduces the contrastive disentangled generative adversarial network (ConDisGAN), a groundbreaking deep architecture designed to overcome challenges in recognition systems, including factors like facial expressions, poses, and lighting variations. ConDisGAN utilizes contrastive learning to effectively capture intrinsic data variations, providing precise control over facial attributes such as neutral expressions, frontal poses, and adaptable illumination. Incorporating 3D priors and an imitation learning algorithm enhances factor disentanglement in the latent space, resulting in substantial improvements compared to existing methods. Rigorous qualitative and quantitative evaluations showcase ConDisGAN’s effectiveness. Visual inspection affirms the high fidelity and diversity of generated facial images. Quantitative comparisons, utilizing established metrics like inception score (IS) and Fréchet inception distance (FID), underscore its superiority over state-of-the-art models. The integration of the mean squared error (MSE) loss function further elevates the quality of the generated images. Collectively, these results underscore ConDisGAN’s potential to advance facial image synthesis technology, enabling the generation of high-quality facial images with precise attribute control.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Facial attribute manipulation using the contrastive disentangled generative adversarial network framework with 3D priors

  • Wiem Grina,
  • Ali Douik

摘要

Facial image synthesis is a pivotal research domain with wide-ranging applications in facial recognition, virtual reality, and the entertainment industry. This paper introduces the contrastive disentangled generative adversarial network (ConDisGAN), a groundbreaking deep architecture designed to overcome challenges in recognition systems, including factors like facial expressions, poses, and lighting variations. ConDisGAN utilizes contrastive learning to effectively capture intrinsic data variations, providing precise control over facial attributes such as neutral expressions, frontal poses, and adaptable illumination. Incorporating 3D priors and an imitation learning algorithm enhances factor disentanglement in the latent space, resulting in substantial improvements compared to existing methods. Rigorous qualitative and quantitative evaluations showcase ConDisGAN’s effectiveness. Visual inspection affirms the high fidelity and diversity of generated facial images. Quantitative comparisons, utilizing established metrics like inception score (IS) and Fréchet inception distance (FID), underscore its superiority over state-of-the-art models. The integration of the mean squared error (MSE) loss function further elevates the quality of the generated images. Collectively, these results underscore ConDisGAN’s potential to advance facial image synthesis technology, enabling the generation of high-quality facial images with precise attribute control.