<p>Traditional methods for sound event detection and localization are difficult to adequately extract discriminative features from complex audio signals, particularly when dealing with signals sharing similar spectral or temporal characteristics. To address this, this paper proposes the AECA-SELD model based on inception depth-wise convolution and adaptive efficient channel attention to improve sound event detection and localization performance. The inception depth-wise convolution blocks captured the deep feature information contained in multi-channel audio signals by decomposing the large kernel depth-wise convolution into multiple small kernel branches. The adaptive efficient channel attention module optimized the feature expression ability of the model and enhanced the flexibility and adaptability of the model feature selection by fusing the local information-enhanced adaptive efficient channel attention and the mixed-scale convolutional feed-forward network. Additionally, the model integrates contextual information across the time–frequency dimensions of sound signals using BiGRU and MHA, enhancing the model’s robustness. Experimental results on the Sony-TAu Reality Spatial Soundscapes 2022 dataset and its extended synthetic dataset show that the <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\({ER}_{20^\circ }\)</EquationSource> </InlineEquation> and the <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\({LE}_{CD}\)</EquationSource> </InlineEquation> of the AECA-SELD model are reduced to 0.66 and 21.2°, respectively, and the <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\({F}_{20^\circ }\)</EquationSource> </InlineEquation> and the <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\({LR}_{CD}\)</EquationSource> </InlineEquation> are increased to 34.9% and 52.5%, respectively, and the comprehensive SELD <InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(Score\)</EquationSource> </InlineEquation> is reduced to 0.48, which outperforms the other compared models. The code is available at <a href="https://github.com/Qiu-sy/AECA-SELD.git">https://github.com/Qiu-sy/AECA-SELD.git</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Sound Event Localization and Detection Research Based on Inception Depth-Wise Convolution and Adaptive Efficient Channel Attention Block

  • Shuyang Qiu,
  • Min Guo,
  • Miao Ma

摘要

Traditional methods for sound event detection and localization are difficult to adequately extract discriminative features from complex audio signals, particularly when dealing with signals sharing similar spectral or temporal characteristics. To address this, this paper proposes the AECA-SELD model based on inception depth-wise convolution and adaptive efficient channel attention to improve sound event detection and localization performance. The inception depth-wise convolution blocks captured the deep feature information contained in multi-channel audio signals by decomposing the large kernel depth-wise convolution into multiple small kernel branches. The adaptive efficient channel attention module optimized the feature expression ability of the model and enhanced the flexibility and adaptability of the model feature selection by fusing the local information-enhanced adaptive efficient channel attention and the mixed-scale convolutional feed-forward network. Additionally, the model integrates contextual information across the time–frequency dimensions of sound signals using BiGRU and MHA, enhancing the model’s robustness. Experimental results on the Sony-TAu Reality Spatial Soundscapes 2022 dataset and its extended synthetic dataset show that the \({ER}_{20^\circ }\) and the \({LE}_{CD}\) of the AECA-SELD model are reduced to 0.66 and 21.2°, respectively, and the \({F}_{20^\circ }\) and the \({LR}_{CD}\) are increased to 34.9% and 52.5%, respectively, and the comprehensive SELD \(Score\) is reduced to 0.48, which outperforms the other compared models. The code is available at https://github.com/Qiu-sy/AECA-SELD.git.