Self-supervised Sound Source Localization for UAVs Using GCC-PHAT in Low SNR Environments
摘要
Recently, utilizing microphone arrays on unmanned aerial vehicle (UAV) platforms for accurate sound source localization has become a critical research direction in machine hearing. However, reliably localizing sound sources in low signal-to-noise ratio (SNR) environments remains highly challenging due to the high-energy narrowband noise generated by UAV propellers and motors. To address these challenges, we propose a self-supervised sound source localization for UAVs in low SNR environments (S3LOC). S3LOC leverages generalized cross-correlation with phase transform (GCC-PHAT) features extracted from microphone arrays and reformulates localization as a grid-based classification task, effectively capturing phase information while mitigating the impact of strong noise interference. The framework incorporates a self-supervised pre-training strategy that combines frequency masking, contrastive learning, and masked reconstruction, enhancing feature robustness and adaptability. Extensive experiments conducted on both simulated and real-world UAV noise datasets demonstrate that S3LOC consistently outperforms existing supervised methods, achieving high localization accuracy even in extremely low SNR scenarios. These results underscore the effectiveness and adaptability of the proposed framework, offering a promising solution for UAV auditory perception and speech processing applications.