Achieving high-quality rendering and real-time generation are pivotal challenges in talking head synthesis. While existing NeRF and 3D Gaussian Splatting (3DGS) based methods deliver rapid, high-fidelity results, they primarily emphasize lip-sync precision, often overlooking nuanced facial expressions and fine-grained detail preservation. To bridge this gap, we present EmoGaussian, an innovative framework that harmonizes accurate lip-sync with natural emotional expressions for real-time talking head generation. Our approach introduces facial action units (AUs) as emotional descriptors, pioneering their integration with 3DGS technology. This enables audio-synchronized dynamic 3D model deformations, ensuring precise alignment of both audio and emotional expressions. Furthermore, we develop an Emotional Enhancement Module that computes attention between decoupled multi-dimensional facial features and real-image-derived facial keypoints, enriching facial detail restoration. Comprehensive experiments demonstrate that EmoGaussian surpasses state-of-the-art methods in lip-sync accuracy, rendering speed, and emotional expressiveness, generating highly realistic and detailed talking head videos.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EmoGaussian High-Fidelity Emotional Talking Head Generation with 3D Gaussian Splatting

  • Tingting Zhang,
  • Zhen Xiao,
  • Jinlin Guo,
  • Xueliang Liu

摘要

Achieving high-quality rendering and real-time generation are pivotal challenges in talking head synthesis. While existing NeRF and 3D Gaussian Splatting (3DGS) based methods deliver rapid, high-fidelity results, they primarily emphasize lip-sync precision, often overlooking nuanced facial expressions and fine-grained detail preservation. To bridge this gap, we present EmoGaussian, an innovative framework that harmonizes accurate lip-sync with natural emotional expressions for real-time talking head generation. Our approach introduces facial action units (AUs) as emotional descriptors, pioneering their integration with 3DGS technology. This enables audio-synchronized dynamic 3D model deformations, ensuring precise alignment of both audio and emotional expressions. Furthermore, we develop an Emotional Enhancement Module that computes attention between decoupled multi-dimensional facial features and real-image-derived facial keypoints, enriching facial detail restoration. Comprehensive experiments demonstrate that EmoGaussian surpasses state-of-the-art methods in lip-sync accuracy, rendering speed, and emotional expressiveness, generating highly realistic and detailed talking head videos.