<p>Gaze estimation plays a pivotal role in the realm of human–machine interaction. Although recent methods perform well in estimating gaze direction by simply fusing the input of single facial image and cropped ocular images, a significant gap remains in the exploration of the correlation and complementarity between these global and local information. To this end, we undertake a systematic attempt to study the role and usage of head pose in gaze estimation, and introduce MixGaze framework, a multi-branch transformer structure with delicately designed mix-attention mechanism, and dual supervision for concurrent head pose prediction and gaze estimation. By explicitly modeling the head pose, decoupling the head and gaze direction at the feature level, and fusing the feature from two parts, MixGaze adeptly steers the model towards the optimal exploitation and discernment of the distinct information encapsulated within facial and ocular inputs. Our work identifies the problems of existing datasets that might hinder previous potential similar explorations, and achieves remarkably competitive performance across three gaze estimation benchmarks. Rigorous ablation analyses substantiate the soundness of our architectural design choices and the generalizable nature of the extracted feature representations, which indicates its potential in practical applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mixgaze: a dually supervised mixed attention network for gaze estimation

  • Ziyang Wu,
  • Yin Lin,
  • Hu Cheng,
  • Caihua Kong,
  • Wengang Zhou,
  • Houqiang Li

摘要

Gaze estimation plays a pivotal role in the realm of human–machine interaction. Although recent methods perform well in estimating gaze direction by simply fusing the input of single facial image and cropped ocular images, a significant gap remains in the exploration of the correlation and complementarity between these global and local information. To this end, we undertake a systematic attempt to study the role and usage of head pose in gaze estimation, and introduce MixGaze framework, a multi-branch transformer structure with delicately designed mix-attention mechanism, and dual supervision for concurrent head pose prediction and gaze estimation. By explicitly modeling the head pose, decoupling the head and gaze direction at the feature level, and fusing the feature from two parts, MixGaze adeptly steers the model towards the optimal exploitation and discernment of the distinct information encapsulated within facial and ocular inputs. Our work identifies the problems of existing datasets that might hinder previous potential similar explorations, and achieves remarkably competitive performance across three gaze estimation benchmarks. Rigorous ablation analyses substantiate the soundness of our architectural design choices and the generalizable nature of the extracted feature representations, which indicates its potential in practical applications.