Recent advancements in artificial intelligence have significantly propelled speech-driven 3D facial animation technologies. However, the existing facial animation generation methods often overlook the critical interplay between latent emotional cues in audio signals and their manifestation in facial dynamics, consequently restricting the fidelity of generated animations. To bridge this gap, we present EC Speaker, a novel conditional diffusion framework that explicitly models emotional speech characteristics to synthesize realistic facial animations. We design a novel emotion-aware extractor that enables effective extraction of latent emotional features from raw audio through the perceptual fusion of time-frequency domain features. In addition, we introduce an emotion-constrained loss function (EmoLoss) to regulate the emotional expression of generated facial animations. Extensive experimental results on the VOCASET dataset demonstrate that the proposed model achieves higher accuracy in facial animation generation compared to other state-of-the-art methods, reducing vertex error and lip error by 41.28% and 75.84%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EC Speaker: Speech-Driven 3D Facial Animation with Latent Emotion Constraints

  • Chen Wang,
  • Wenyuan Ying,
  • Tianyang Dong

摘要

Recent advancements in artificial intelligence have significantly propelled speech-driven 3D facial animation technologies. However, the existing facial animation generation methods often overlook the critical interplay between latent emotional cues in audio signals and their manifestation in facial dynamics, consequently restricting the fidelity of generated animations. To bridge this gap, we present EC Speaker, a novel conditional diffusion framework that explicitly models emotional speech characteristics to synthesize realistic facial animations. We design a novel emotion-aware extractor that enables effective extraction of latent emotional features from raw audio through the perceptual fusion of time-frequency domain features. In addition, we introduce an emotion-constrained loss function (EmoLoss) to regulate the emotional expression of generated facial animations. Extensive experimental results on the VOCASET dataset demonstrate that the proposed model achieves higher accuracy in facial animation generation compared to other state-of-the-art methods, reducing vertex error and lip error by 41.28% and 75.84%.