ARG-Net: Gaze Estimation Based on Adversarial Learning and Learnable Networks
摘要
Gaze estimation tasks have widespread applications in fields such as human-computer interaction, virtual reality, and driver monitoring. However, they still face several challenges and difficulties in practical scenarios. Issues such as individual differences, occlusion, lighting variations, dynamic environments, and difficulties in data annotation continue to pose challenges for this task in real-world applications. Therefore, this paper proposes an innovative gaze estimation model—ARG-Net—based on cross-domain gaze estimation tasks. This model first designs a facial feature adversarial reconstruction module to remove unnecessary background noise and retain the key information related to gaze direction. Then, after processing the features, the image is passed through a multi-head self-attention mechanism to further enhance the ability to capture complex contextual information, effectively extracting long-range dependencies from facial images. This helps the model accurately predict gaze direction even in complex environments. The network’s non-linear representation ability is enhanced through learnable activation functions, improving its ability to perceive subtle eye movements and handle complex gaze features. Experimental results show that ARG-Net performs excellently in cross-domain gaze estimation tasks on the Gaze360 and MPiiGaze datasets, with gaze estimation errors of 9.84° and 6.53°, respectively, lower than the recent cross-domain gaze estimation errors in these two datasets. Compared with other existing models, ARG-Net achieves better accuracy and robustness, performing well under occlusion, lighting changes, and complex scenarios. It has broad practical application potential and can provide more accurate gaze detection solutions for intelligent interaction, security monitoring, and other fields.