JoyLive: Efficient Audio-Driven Portrait Animation by 3D Implict Keypoints
摘要
Audio-driven portrait animation aims at generating realistic and natural talking head videos from audio inputs and static portrait images. While recent stable diffusion-based methods have achieved impressive visual quality, they often suffer from high computational costs and low inference efficiency. In this paper, we propose JoyLive, an intermediate representation-based approach that leverages implicit keypoints as motion representations for audio-driven portrait animation. Specifically, 3D implicit keypoints model the overall understanding of the face, containing fully decoupled semantic and geometric information, which can be seamlessly integrated into our carefully designed lightweight diffusion transformer architecture for more flexible control of facial expressions and head movements. To further preserve identity consistency, we introduce a conditioning mechanism that integrates the canonical keypoints of the first frame through a cross-attention module during the diffusion process. Additionally, we incorporate image perceptual loss and head pose loss during training to further enhance image generation quality and vivid head movements. Experimental results on the HDTF dataset demonstrate that our method achieves competitive performance compared to existing approaches, while not only significantly enhancing generation efficiency and video quality, but also resulting in richer facial expressions and more natural head movements.