错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Audio-Visual Sound Event Localization and Detection Based on CRNN Using Depth-Wise Separable Convolution

  • Yi Wang,
  • Hongqing Liu,
  • Yu Zhao,
  • Yi Zhou

摘要

Sound event localization and detection (SELD) focuses on the simultaneous detection of various sound events along with their sptial and temporal localization. Recent work shows that audio-visual fusion methods, rarely involved in SELD research, show promising results than the single modality. For this, we proposed an audio and visual signals fusion mechanism for SELD based on convolutional recurrent neural network (CRNN). Object detection and pre-trained model processing on the corresponding image at the start frame of the audio feature sequence is utilized to acquire visual cues passed into models. Compared to traditional convolution, we devise a depth-wise separable convolution block to better learn the relevant information of different sound event categories in audio features. Experimental results on STARSS23 of DCASE (2023) indicate that the introducing of visual cues do improve the SELD performance compared to the audio-only system. The convolution block devised in our proposed work further enhances the model’s performance as it achieves higher SELD score.