EmoGaussian High-Fidelity Emotional Talking Head Generation with 3D Gaussian Splatting
摘要
Achieving high-quality rendering and real-time generation are pivotal challenges in talking head synthesis. While existing NeRF and 3D Gaussian Splatting (3DGS) based methods deliver rapid, high-fidelity results, they primarily emphasize lip-sync precision, often overlooking nuanced facial expressions and fine-grained detail preservation. To bridge this gap, we present EmoGaussian, an innovative framework that harmonizes accurate lip-sync with natural emotional expressions for real-time talking head generation. Our approach introduces facial action units (AUs) as emotional descriptors, pioneering their integration with 3DGS technology. This enables audio-synchronized dynamic 3D model deformations, ensuring precise alignment of both audio and emotional expressions. Furthermore, we develop an Emotional Enhancement Module that computes attention between decoupled multi-dimensional facial features and real-image-derived facial keypoints, enriching facial detail restoration. Comprehensive experiments demonstrate that EmoGaussian surpasses state-of-the-art methods in lip-sync accuracy, rendering speed, and emotional expressiveness, generating highly realistic and detailed talking head videos.