Mixgaze: a dually supervised mixed attention network for gaze estimation
摘要
Gaze estimation plays a pivotal role in the realm of human–machine interaction. Although recent methods perform well in estimating gaze direction by simply fusing the input of single facial image and cropped ocular images, a significant gap remains in the exploration of the correlation and complementarity between these global and local information. To this end, we undertake a systematic attempt to study the role and usage of head pose in gaze estimation, and introduce MixGaze framework, a multi-branch transformer structure with delicately designed mix-attention mechanism, and dual supervision for concurrent head pose prediction and gaze estimation. By explicitly modeling the head pose, decoupling the head and gaze direction at the feature level, and fusing the feature from two parts, MixGaze adeptly steers the model towards the optimal exploitation and discernment of the distinct information encapsulated within facial and ocular inputs. Our work identifies the problems of existing datasets that might hinder previous potential similar explorations, and achieves remarkably competitive performance across three gaze estimation benchmarks. Rigorous ablation analyses substantiate the soundness of our architectural design choices and the generalizable nature of the extracted feature representations, which indicates its potential in practical applications.