GaPTalk: precision-controlled 3D Gaussian rendering for personalized talking-head synthesis
摘要
Photorealistic audio-driven talking-head synthesis is pivotal for immersive human–computer interaction, yet faces challenges in semantic alignment, personalization, and efficiency. This paper introduces GaPTalk, a novel framework leveraging precision-controlled 3D Gaussians for personalized talking face generation. GaPTalk integrates a contextual audio-to-expression encoding module, a precision-controlled 3D Gaussian rendering module, and a global–local inpainting module for seamless head-body reenactment. Extensive evaluations demonstrate that GaPTalk outperforms state-of-the-art methods in visual quality, lip-sync accuracy, and computational efficiency, achieving the fastest training time (< 1 h) and real-time inference. Here, we show that GaPTalk provides a robust and efficient solution for high-quality personalized digital human synthesis. Code and video demos are available at https://github.com/davisleelx/GaPTalk.