We present LW-GeneFace, a lightweight and high-fidelity model for generalized audio-driven facial animation, in this paper. We develop this model by reducing the size while maintaining the synthetic quality of GeneFace, an audio-driven facial animation model known for its high fidelity and generalization capabilities. Specifically, we compress the first and the third stages of GeneFace as they dominate the model size. In the first stage, we propose a lightweight version of the WaveNet-based network inspired by MobileNetV3 and DP-block. It utilizes depthwise separable convolution and dual-path feature extraction to compress the network while maintaining effective feature extraction. The shared network structure in the dual-path feature extraction further reduces model complexity and improves training efficiency. In the third stage, we generate realistic 3D renderings at reduced model size by introducing novelties in RAD-NeRF. Technically, we reduce the hash table sizes in the grid-based encoding modules, as well as present a lightweight bottleneck MLP architecture to increase the non-linearity of the model. Experimental results demonstrate that LW-GeneFace achieves state-of-the-art performance with both model size and synthetic quality considered. The source code of LW-GeneFace will be released after acceptance of this paper.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Lightweight and High-Fidelity Model for Generalized Audio-Driven 3D Talking Face Synthesis

  • Shunce Liu,
  • Yuwei Zhong,
  • Huixuan Wang,
  • Jingliang Peng

摘要

We present LW-GeneFace, a lightweight and high-fidelity model for generalized audio-driven facial animation, in this paper. We develop this model by reducing the size while maintaining the synthetic quality of GeneFace, an audio-driven facial animation model known for its high fidelity and generalization capabilities. Specifically, we compress the first and the third stages of GeneFace as they dominate the model size. In the first stage, we propose a lightweight version of the WaveNet-based network inspired by MobileNetV3 and DP-block. It utilizes depthwise separable convolution and dual-path feature extraction to compress the network while maintaining effective feature extraction. The shared network structure in the dual-path feature extraction further reduces model complexity and improves training efficiency. In the third stage, we generate realistic 3D renderings at reduced model size by introducing novelties in RAD-NeRF. Technically, we reduce the hash table sizes in the grid-based encoding modules, as well as present a lightweight bottleneck MLP architecture to increase the non-linearity of the model. Experimental results demonstrate that LW-GeneFace achieves state-of-the-art performance with both model size and synthetic quality considered. The source code of LW-GeneFace will be released after acceptance of this paper.