The existing audio-visual wake-up word spotting (AVWWS) methods assume that the audio signal has been aligned with the lip movement video signal of a specific speaker in noisy environments, and are mainly applicable for scenarios with only a single speaker. However, in complex scenarios, there may be multiple people showing up in the video facing the camera simultaneously, and more than one person may be speaking at the same time. Wake-up word spotting in noisy and multi-person scenarios remains relatively under-explored. In this paper, we first propose a Wake-up Word Active Speaker Detection Model (WWASD) to recognize the face that is speaking the wake-up word. Based on the model, we propose two approaches, namely Two-stage detection and Three-stage detection, for audio-visual wake-up word spotting in noisy and multi-person scenarios. We compare the approaches from the perspectives of performance and computational complexity on MISP2021-AVWWS corpus. The best Two-stage detection approach, which contains WWASD and audio-visual wake-up word spotting model, achieves comparable performance against the systems with oracle visual speaker bounding boxes. Three-stage detection, which adds an audio-based single-modality wake-up word model as a front end greatly reduces the computational cost.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Audio-Visual Wake-up Word Spotting Under Noisy and Multi-person Scenarios

  • Cancan Li,
  • Fei Su,
  • Juan Liu

摘要

The existing audio-visual wake-up word spotting (AVWWS) methods assume that the audio signal has been aligned with the lip movement video signal of a specific speaker in noisy environments, and are mainly applicable for scenarios with only a single speaker. However, in complex scenarios, there may be multiple people showing up in the video facing the camera simultaneously, and more than one person may be speaking at the same time. Wake-up word spotting in noisy and multi-person scenarios remains relatively under-explored. In this paper, we first propose a Wake-up Word Active Speaker Detection Model (WWASD) to recognize the face that is speaking the wake-up word. Based on the model, we propose two approaches, namely Two-stage detection and Three-stage detection, for audio-visual wake-up word spotting in noisy and multi-person scenarios. We compare the approaches from the perspectives of performance and computational complexity on MISP2021-AVWWS corpus. The best Two-stage detection approach, which contains WWASD and audio-visual wake-up word spotting model, achieves comparable performance against the systems with oracle visual speaker bounding boxes. Three-stage detection, which adds an audio-based single-modality wake-up word model as a front end greatly reduces the computational cost.