A Dual-Path Approach for Gaze Following in Fisheye Meeting Scenes
摘要
Gaze following plays a crucial role in scene comprehension tasks, as it captures users’ visual information from their facial and eye movements, thereby predicting their gaze positions. This technique finds its application in various domains such as human-computer interaction and medical diagnosis. In the domain of multi-party meeting scenes, some studies have utilized fisheye cameras to capture the entire meeting scene. In this work, we focus on gaze following methods that utilize fisheye images for meeting scenes and collect the GazeMeeting dataset that contains 31,915 fisheye samples. We also propose a dual-path feature fusing model for gaze following, which fuses the learned features in the planar and spherical domains by introducing spherical convolutions. The dual-pathway model can learn the distortion information of different positions from scene images, achieving a normalized L2 distance of 0.0657 on our self-built GazeMeeting dataset. This result represents a 22.80% improvement over the current state-of-the-art methods. Additionally, our proposed model achieves a normalized L2 distance of 0.1326 on GazeFollow dataset, outperforming the current state-of-the-art methods by 3.35%.