<p>As a fundamental step in array signal processing, accurate direction-of-arrival (DOA) estimation is crucial for speaker localization using microphone arrays. Noise, reverberation, and an unknown number of sources in realistic environments pose significant challenges, making the extraction of discriminative representations a key step in DOA estimation. These representations need to reduce the influence of redundant information unrelated to localization, yet recent methods have largely overlooked this important characteristic. To address these issues, we propose an end-to-end feature integration and discriminative learning network (FID-Net) for multi-source DOA estimation. Specifically, our approach consists of three stages: the feature integration stage, the discriminative learning stage, and the temporal modeling stage. In the feature integration stage, we aim to capture multi-scale spatial information that is critical for localization. In the discriminative learning stage, we introduce a discriminative representation learning strategy and design a mutual information-based loss to guide the network to better capture the differences among diverse features. The discriminative features are further utilized in the temporal modeling stage to enhance the global contextual representation. Experimental results on both simulated and real-world datasets demonstrate the superior performance of the proposed method compared with other advanced methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning discriminative representations from integrated features for DOA estimation

  • Qi You,
  • Qinghua Huang

摘要

As a fundamental step in array signal processing, accurate direction-of-arrival (DOA) estimation is crucial for speaker localization using microphone arrays. Noise, reverberation, and an unknown number of sources in realistic environments pose significant challenges, making the extraction of discriminative representations a key step in DOA estimation. These representations need to reduce the influence of redundant information unrelated to localization, yet recent methods have largely overlooked this important characteristic. To address these issues, we propose an end-to-end feature integration and discriminative learning network (FID-Net) for multi-source DOA estimation. Specifically, our approach consists of three stages: the feature integration stage, the discriminative learning stage, and the temporal modeling stage. In the feature integration stage, we aim to capture multi-scale spatial information that is critical for localization. In the discriminative learning stage, we introduce a discriminative representation learning strategy and design a mutual information-based loss to guide the network to better capture the differences among diverse features. The discriminative features are further utilized in the temporal modeling stage to enhance the global contextual representation. Experimental results on both simulated and real-world datasets demonstrate the superior performance of the proposed method compared with other advanced methods.