A Sound Event Localization and Detection Research Based on Inception Depth-Wise Convolution and Adaptive Efficient Channel Attention Block
摘要
Traditional methods for sound event detection and localization are difficult to adequately extract discriminative features from complex audio signals, particularly when dealing with signals sharing similar spectral or temporal characteristics. To address this, this paper proposes the AECA-SELD model based on inception depth-wise convolution and adaptive efficient channel attention to improve sound event detection and localization performance. The inception depth-wise convolution blocks captured the deep feature information contained in multi-channel audio signals by decomposing the large kernel depth-wise convolution into multiple small kernel branches. The adaptive efficient channel attention module optimized the feature expression ability of the model and enhanced the flexibility and adaptability of the model feature selection by fusing the local information-enhanced adaptive efficient channel attention and the mixed-scale convolutional feed-forward network. Additionally, the model integrates contextual information across the time–frequency dimensions of sound signals using BiGRU and MHA, enhancing the model’s robustness. Experimental results on the Sony-TAu Reality Spatial Soundscapes 2022 dataset and its extended synthetic dataset show that the