错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-Modal Active Speaker Detection Algorithm in Video and End-To-End Landing Solution

  • Yatao Yang,
  • Siyu Yan

摘要

This research paper delves into the realm of active speaker detection (ASD) in video content with the primary objective of accurately identifying one or more active speakers. In contrast to traditional speech recognition, ASD focuses on verifying the sources of sound rather than the specific content of speech. Unlike speaker diarization, which categorizes speech segments by distinct speakers, ASD precisely identifies the active speaker in scenarios involving multiple speakers. The proposed OneStage architecture adopts a self-attention-based network to dynamically integrate cross-modal information. Notably, the incorporation of differential images serves to diminish the reliance on visual information. To alleviate the burden of annotation, a novel random ratio mask annotation and audio-visual asynchronous data augmentation are introduced, significantly reducing the annotation workload and enhancing the generalization of the model. Furthermore, the paper introduces the OnePass architecture tailored for lightweight real-time speaker detection on small terminal devices. With a modest 5 million parameters, 0.05 GFLOPs computation, and an impressive accuracy rate of 89.5%, the OnePass architecture achieves real-time speaker detection. Meanwhile, the OneStage architecture demonstrates exceptional performance with a 95.98% area under the curve (AUC) on the test set. These findings showcase the efficacy of the proposed architectures in advancing the field of active speaker detection in videos.