Localization of Multiple Sound Sources Based on Deep Learning Using a Microphone Array and Acoustic Images
摘要
A deep learning method based on functional beamforming images and a lightweight densely connected fully convolutional neural network is proposed to quickly and precisely estimate the number, positions, and strengths of multiple sound sources. Functional beamforming is employed to obtain the sound source distribution images as inputs of the network. Then, a target function is defined to generate target images as the ground truth labels for training the network. Finally, a lightweight densely connected fully convolutional neural network with encoder-decoder structure is established for the localization task of multiple sound sources. One encoder network is used for down-sampling the functional beamforming images and capturing the context information and spatial features. One decoder network implements up-sampling operations for transforming the feature images into the predicted target images. The feature images generated by diverse layers in the dense block are concatenated through dense connections to enhance the flow of information and gradients within the network while improving its ability to reuse features. One to three monopoles are distributed randomly on the plane of 1 m × 1 m with 11 random strengths to generate various sound fields at frequencies of 1 kHz, 2 kHz, and 6 kHz. The quantitative and qualitative comparison results prove that the proposed method can predict the number, positions, and strengths of multiple sound sources well, and its localization accuracy and computation efficiency are better than those of the conventional beamforming-based fully convolutional network. The stronger stability and robustness of the proposed method are also testified under unseen acoustic conditions. An experiment conducted in a semi-anechoic chamber with two loudspeakers is employed to further validate the effectiveness of the proposed method. On the simulated dataset, the proposed model achieves a high localization accuracy of 99.6%, with distance and strength errors of only 0.00291 m and 0.51 dB, respectively. Additionally, the proposed model trained on the simulated dataset can accurately identify two loudspeakers with a center distance of 0.24 m in a semi-anechoic chamber within a broad range of frequencies.