<p>Talking face generation has seen remarkable progress, yet achieving precise lip synchronization, natural expression dynamics, and high visual realism remains challenging. Many existing methods suffer from abnormalities with artifacts in the mouth region and inconsistencies in the speaker’s identity across frames. To address these limitations, we propose a latent-space framework that refines mouth features before synthesizing the full face, enhancing articulation accuracy and perceptual fidelity. Our approach incorporates a Variational Autoencoder (VAE)-based latent fusion strategy to balance computational space with structural consistency while preserving dynamic facial expressions. Additionally, a reference-guided mechanism improves mouth-region synthesis, reducing artifacts and enhancing identity retention. Experiments on the MEAD and CREMA-D datasets demonstrate that our method achieves competitive performance compared to state-of-the-art techniques, offering improved lip-sync accuracy, expression coherence, and visual quality. The implementation and additional details can be found at <a href="https://github.com/cxnam-vnuhcmus/PMFS.git">https://github.com/cxnam-vnuhcmus/PMFS.git</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Pmfs: Progressive mouth-to-face synthesis for realistic talking face generation

  • Xuan-Nam Cao,
  • Nhat-Tan Vo,
  • Minh-Triet Tran

摘要

Talking face generation has seen remarkable progress, yet achieving precise lip synchronization, natural expression dynamics, and high visual realism remains challenging. Many existing methods suffer from abnormalities with artifacts in the mouth region and inconsistencies in the speaker’s identity across frames. To address these limitations, we propose a latent-space framework that refines mouth features before synthesizing the full face, enhancing articulation accuracy and perceptual fidelity. Our approach incorporates a Variational Autoencoder (VAE)-based latent fusion strategy to balance computational space with structural consistency while preserving dynamic facial expressions. Additionally, a reference-guided mechanism improves mouth-region synthesis, reducing artifacts and enhancing identity retention. Experiments on the MEAD and CREMA-D datasets demonstrate that our method achieves competitive performance compared to state-of-the-art techniques, offering improved lip-sync accuracy, expression coherence, and visual quality. The implementation and additional details can be found at https://github.com/cxnam-vnuhcmus/PMFS.git.