Learning Collaborative Reinforcement Attention for 3D Face Reconstruction and Dense Alignment
摘要
3D face reconstruction from monocular outdoor images has long been a challenging problem. Traditional attention network methods that directly regress parameters may suffer from inadequate learning of discriminative features. In this paper, we propose a method called collaborative reinforcement attention module (CRAM). CRAM comprises three major modules: the perception module (PM), the channel selection module (CSM), and the multi-level feature interaction module (MFIM). CRAM leverages contextual information to simultaneously focus on multiple prominent features in facial photos. It employs multi-level and multi-angle feature extraction and fusion techniques to adaptively learn the relationship between facial regions and key feature points. This results in enhanced accuracy in 3D face reconstruction and meticulous dense alignment. Furthermore, to enhance the model’s generalization performance, we introduce a regional noise injection and image composition module (RNICM) as a preprocessing step for sample data which help capture more local details and handle occluded faces, particularly under significant head rotations. Extensive experiments conducted on the AFLW2000-3D and AFLW datasets validate the effectiveness of the proposed approach.