Existing NeRF-based head avatar reconstruction methods utilize expression coefficients as driving signals. Despite significant advancements, they fail to accurately capture facial feature deformations under complex expression changes. To address this issue, we integrate prior information from public facial feature dictionaries with expression coefficients as driving signals, and employ a region attention mechanism to more accurately capture facial deformations. First, when reconstructing a single face, we extract the facial features from a single image and obtain prior information on these features from a public facial feature dictionary. This prior information is integrated with expression coefficients as driving signals to more accurately drive the deformation of facial features. Second, we introduce a region attention mechanism that learns the explicit relationship between local spatial regions and driving signals during training. This allows for differential driving effects on various facial regions with the same driving signal, achieving more precise local motion modeling. Experimental results show that our method can precisely capture subtle deformations across all facial regions and outperforms state-of-the-art methods in both qualitative and quantitative aspects.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DictAvatar: Expressive Facial Avatar Reconstruction with Facial Feature Dictionary

  • Zuyin Wu,
  • Zixuan Guo,
  • Yifan Xie,
  • Tiffany He,
  • Fei Ma,
  • Fei Yu

摘要

Existing NeRF-based head avatar reconstruction methods utilize expression coefficients as driving signals. Despite significant advancements, they fail to accurately capture facial feature deformations under complex expression changes. To address this issue, we integrate prior information from public facial feature dictionaries with expression coefficients as driving signals, and employ a region attention mechanism to more accurately capture facial deformations. First, when reconstructing a single face, we extract the facial features from a single image and obtain prior information on these features from a public facial feature dictionary. This prior information is integrated with expression coefficients as driving signals to more accurately drive the deformation of facial features. Second, we introduce a region attention mechanism that learns the explicit relationship between local spatial regions and driving signals during training. This allows for differential driving effects on various facial regions with the same driving signal, achieving more precise local motion modeling. Experimental results show that our method can precisely capture subtle deformations across all facial regions and outperforms state-of-the-art methods in both qualitative and quantitative aspects.