FAFNet: Feature Adaptive Fusion Network for Robust Appearance-Based Gaze Estimation Under Extreme Head Poses
摘要
As deep learning has rapidly progressed in recent years, app-earance-based gaze estimation has attracted considerable attention. However, gaze estimation remains challenging because of limited resolution, low lighting, and the inability to acquire eye images as a result of extreme head pose. Mainstream approaches focus primarily on efficient feature extraction while ignoring feature fusion. Herein, we propose a gaze-estimating network, FAFNet, to address the issues of extreme head pose. The proposed FAFNet incorporates all convolutional layer features based on channel attention, reducing error accumulation due to missing information. Moreover, we develop an adaptive feature fusion block based on cross-attention to merge eye and facial features, through alternating queries of global and local features, thereby addressing the problems associated with an extreme head pose. Extensive tests were conducted on three datasets: ETH-XGaze, MPIIFaceGaze, and EYEDIAP. Results show that the proposed FAFNet surpasses state-of-the-art approaches ETH-XGaze dataset and achieves performance comparable to the leading methods on the MPIIFaceGaze and EYEDIAP datasets. The performance of our network is further validated by ablation experiments.