The traditional 3D morphable model (3DMM) techniques generally regress model coefficients directly, neglecting the critical 2D spatial and semantic edge information. To address this limitation, we propose a multi-branch attention network (MBAN) designed to reconstruct 3D faces from monocular outdoor images, effectively mitigating edge feature loss. Our approach includes preprocessing through image cropping and adaptive edge feature learning using an edge facial feature extraction module. Additionally, we introduce a global feature extraction module based on an enhanced deformator, which enables the interaction between semantic and spatial information of global features for fine-grained representation. This study presents three attention mechanisms: the edge attention mechanism (EAM), the global attention mechanism (GAM), and the local attention mechanism (LAM). These mechanisms establish correlations among edge, global, and local features, enhancing feature extraction at various levels and facilitating multi-level interactive feature learning. This approach addresses the challenges of 3D face reconstruction under varying poses and complex environments, thereby improving the robustness of the reconstruction process. Extensive experiments on the AFLW2000-3D and AFLW datasets validate the effectiveness of our proposed method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning Multi-Branch Attention Networks for 3D Face Reconstruction

  • Lei Ma,
  • Zhengwei Yang,
  • Yange Wang,
  • Xiangzheng Li

摘要

The traditional 3D morphable model (3DMM) techniques generally regress model coefficients directly, neglecting the critical 2D spatial and semantic edge information. To address this limitation, we propose a multi-branch attention network (MBAN) designed to reconstruct 3D faces from monocular outdoor images, effectively mitigating edge feature loss. Our approach includes preprocessing through image cropping and adaptive edge feature learning using an edge facial feature extraction module. Additionally, we introduce a global feature extraction module based on an enhanced deformator, which enables the interaction between semantic and spatial information of global features for fine-grained representation. This study presents three attention mechanisms: the edge attention mechanism (EAM), the global attention mechanism (GAM), and the local attention mechanism (LAM). These mechanisms establish correlations among edge, global, and local features, enhancing feature extraction at various levels and facilitating multi-level interactive feature learning. This approach addresses the challenges of 3D face reconstruction under varying poses and complex environments, thereby improving the robustness of the reconstruction process. Extensive experiments on the AFLW2000-3D and AFLW datasets validate the effectiveness of our proposed method.