A Lightweight and High-Fidelity Model for Generalized Audio-Driven 3D Talking Face Synthesis
摘要
We present LW-GeneFace, a lightweight and high-fidelity model for generalized audio-driven facial animation, in this paper. We develop this model by reducing the size while maintaining the synthetic quality of GeneFace, an audio-driven facial animation model known for its high fidelity and generalization capabilities. Specifically, we compress the first and the third stages of GeneFace as they dominate the model size. In the first stage, we propose a lightweight version of the WaveNet-based network inspired by MobileNetV3 and DP-block. It utilizes depthwise separable convolution and dual-path feature extraction to compress the network while maintaining effective feature extraction. The shared network structure in the dual-path feature extraction further reduces model complexity and improves training efficiency. In the third stage, we generate realistic 3D renderings at reduced model size by introducing novelties in RAD-NeRF. Technically, we reduce the hash table sizes in the grid-based encoding modules, as well as present a lightweight bottleneck MLP architecture to increase the non-linearity of the model. Experimental results demonstrate that LW-GeneFace achieves state-of-the-art performance with both model size and synthetic quality considered. The source code of LW-GeneFace will be released after acceptance of this paper.