Pmfs: Progressive mouth-to-face synthesis for realistic talking face generation
摘要
Talking face generation has seen remarkable progress, yet achieving precise lip synchronization, natural expression dynamics, and high visual realism remains challenging. Many existing methods suffer from abnormalities with artifacts in the mouth region and inconsistencies in the speaker’s identity across frames. To address these limitations, we propose a latent-space framework that refines mouth features before synthesizing the full face, enhancing articulation accuracy and perceptual fidelity. Our approach incorporates a Variational Autoencoder (VAE)-based latent fusion strategy to balance computational space with structural consistency while preserving dynamic facial expressions. Additionally, a reference-guided mechanism improves mouth-region synthesis, reducing artifacts and enhancing identity retention. Experiments on the MEAD and CREMA-D datasets demonstrate that our method achieves competitive performance compared to state-of-the-art techniques, offering improved lip-sync accuracy, expression coherence, and visual quality. The implementation and additional details can be found at https://github.com/cxnam-vnuhcmus/PMFS.git.