Cross-Modal Active Speaker Detection Algorithm in Video and End-To-End Landing Solution
摘要
This research paper delves into the realm of active speaker detection (ASD) in video content with the primary objective of accurately identifying one or more active speakers. In contrast to traditional speech recognition, ASD focuses on verifying the sources of sound rather than the specific content of speech. Unlike speaker diarization, which categorizes speech segments by distinct speakers, ASD precisely identifies the active speaker in scenarios involving multiple speakers. The proposed OneStage architecture adopts a self-attention-based network to dynamically integrate cross-modal information. Notably, the incorporation of differential images serves to diminish the reliance on visual information. To alleviate the burden of annotation, a novel random ratio mask annotation and audio-visual asynchronous data augmentation are introduced, significantly reducing the annotation workload and enhancing the generalization of the model. Furthermore, the paper introduces the OnePass architecture tailored for lightweight real-time speaker detection on small terminal devices. With a modest 5 million parameters, 0.05 GFLOPs computation, and an impressive accuracy rate of 89.5%, the OnePass architecture achieves real-time speaker detection. Meanwhile, the OneStage architecture demonstrates exceptional performance with a 95.98% area under the curve (AUC) on the test set. These findings showcase the efficacy of the proposed architectures in advancing the field of active speaker detection in videos.