PAED: Physics-Anchored and Emotion Disentangled Memory Diffusion for Talking Face Generation
摘要
The ability to generate emotionally expressive audio-based videos has seen significant advancements with end-to-end diffusion models, which have greatly improved video quality. However, these models still struggle with expression inconsistencies and unnatural transitions, particularly in long-duration sequences and complex emotional shifts. To address these limitations, we propose Physics-Anchored and Emotion Disentangled Memory Diffusion (PAED), a hybrid framework that enhances pre-trained memory architectures through two novel modules: (1) a Physics-Guided AU Momentum Regulator, which models facial dynamics as damped harmonic systems to enforce biomechanically valid AU transitions through inertial constraints, and (2) a Cross-Modal Emotion Disentangler, which employs contrastive momentum queues to align audio-visual emotion representations while minimizing mutual information between emotion and articulation features. Experiments show that our model reduces expression inconsistency compared to previous models, offering a more emotionally nuanced user experience. Access to the source code can be found at: https://github.com/YummyBamboo/PAED_MEMO .