Robust Active Speaker Detection in Challenging Environments Using GNN-Fused Multi-modal Cues and Body Language
摘要
Active Speaker Detection (ASD) is a multi-modal task which aims to identify speakers using both audio and visual cues. Existing approaches often focus solely on facial and audio features, overlooking the importance of body language. This motivates us to integrate body movements as an additional cue to enhance ASD performance, especially in challenging conditions such as low brightness or high noise levels. Furthermore, we propose to use graph neural networks (GNNs) to fuse multiple contexts, which is more flexible and effective than existing methods. Experimental results demonstrate the superiority of our method over the state-of-the-art approaches on two popular datasets. Compared to the most competitive counterparts, our method achieves a mean Average Precision (mAP) improvement of 0.6% on AVA-ActiveSpeaker and 4.88% on the challenging Ego4D-ASD dataset.