Blendshape-Based Migratable Speech-Driven 3D Facial Animation with Overlapping Chunking-Transformer
摘要
Speech-driven 3D facial animation has attracted an amount of research and has been widely used in games and virtual reality. Most of the latest state-of-the-art methods employ Transformer-based architecture with good sequence modeling capability. However, most of the animations produced by these methods are limited to specific facial meshes and cannot handle lengthy audio inputs. To tackle these limitations, we leverage the advantage of blendshapes to migrate the generated animations to multiple facial meshes and propose an overlapping chunking strategy that enables the model to support long audio inputs. Also, we design a data calibration approach that can significantly enhance the quality of blendshapes data and make lip movements more natural. Experiments show that our method performs better than the methods predicting vertices, and the animation can be migrated to various meshes.