Modular Joint Training for Speech-Driven 3D Facial Animation
摘要
Speech-driven 3D facial animation is still an intensive field of research, with some persistent challenges. These difficulties arise from the intricate nature of achieving facial realism and the scarcity of audiovisual data. Previous studies have mainly focused on learning phoneme-level features from brief audio segments, which often lead to suboptimal lip movements. To capture the nuances of facial expressions, such as eyebrow-raising or lip curling, our proposed solution builds upon the autoregressive model of the Transformer, which can generate realistic facial movements based on previous frames and introduces a modular face separation model, which can separately control the upper face and lips, to enhance the quality of voice-driven 3D facial animation. This novel modular separation technique divides the facial mesh into two parts: the upper face and the lips, using our uniquely designed mask. Such an approach not only significantly improves facial animation synthesis but also lays the foundation for future research and application in this domain.