Gaze-Driven Active Speaker Detection in Meetings
摘要
Active Speaker Detection (ASD) plays a vital role in scene understanding tasks, aiming to determine whether individuals in a scene are speaking. It has broad applications in areas such as speaker diarization and speaker tracking. Mainstream approaches typically rely on facial images and audio spectrograms to make predictions. In this paper, we explore the ASD task in fisheye meeting scenes and highlight the observation that participants tend to gaze at the active speaker. Based on this insight, we propose a gaze-driven multimodal ASD method. Specifically, we aggregate the gaze directions of all participants to construct a scene-level gaze field. A dedicated feature extraction branch then captures participants’ attention patterns, enhancing ASD accuracy. Furthermore, to address facial distortions caused by fisheye cameras, we employ Deformable Convolution networks (DConv). And we use Time Delay Neural Networks (TDNN) to extract temporal features from audio sequences. We evaluate our method on the FisheyeMeeting dataset, which targets multi-speaker detection in real-world meeting scenes. Experimental results demonstrate that our method significantly improves ASD accuracy in real-world meeting scenes.