Existing acoustic localization methods face challenges in dynamic sound source scenarios due to multipath interference, leading to spatial ambiguity. Although multimodal fusion methods can improve accuracy, they rely on real-time visual data streams, resulting in degraded performance in scenarios with visual failures, such as darkness or occlusion. To this end, we propose the Visual-Guided Decoupled Acoustic Localization Network (ViDAL-Net), which achieves robust localization in environments through cross-modal knowledge distillation and modal decoupling mechanisms. ViDAL-Net leverages visual target detection data as supervisory signals to implicitly encode visual geometric priors into the acoustic feature space, establishing a mapping between acoustic responses and image planes. The network architecture employs joint optimization of coordinate regression and heatmap prediction tasks. Coordinate regression employs a learnable B-spline-based Kernel Approximation Network (KAN) network to achieve superior continuous spatial localization capabilities compared to traditional discrete grid search methods. The heatmap prediction task utilizes an adaptive Gaussian kernel decoder to capture spatial distribution patterns. The dual-task collaborative mechanism effectively balances geometric precision and semantic reasonableness. Furthermore, we introduce a visual-sound decoupling inference paradigm, where the training stage incorporates visual supervisory information, while the inference stage relies solely on acoustic signals for mono-modal localization, thus breaking the traditional dependency on real-time visual data. Experimental results on the AV16.3 dataset demonstrate that ViDAL-Net achieves superior localization performance in visual failure scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Vision-Guided Acoustic Localization with Decoupled Inference for Moving Speakers

  • Yidi Li,
  • Kairan Zhang,
  • Chenxu Yang,
  • Chongwei Yan,
  • Rongshan Gao,
  • Mingliang Dou,
  • Bin Ren

摘要

Existing acoustic localization methods face challenges in dynamic sound source scenarios due to multipath interference, leading to spatial ambiguity. Although multimodal fusion methods can improve accuracy, they rely on real-time visual data streams, resulting in degraded performance in scenarios with visual failures, such as darkness or occlusion. To this end, we propose the Visual-Guided Decoupled Acoustic Localization Network (ViDAL-Net), which achieves robust localization in environments through cross-modal knowledge distillation and modal decoupling mechanisms. ViDAL-Net leverages visual target detection data as supervisory signals to implicitly encode visual geometric priors into the acoustic feature space, establishing a mapping between acoustic responses and image planes. The network architecture employs joint optimization of coordinate regression and heatmap prediction tasks. Coordinate regression employs a learnable B-spline-based Kernel Approximation Network (KAN) network to achieve superior continuous spatial localization capabilities compared to traditional discrete grid search methods. The heatmap prediction task utilizes an adaptive Gaussian kernel decoder to capture spatial distribution patterns. The dual-task collaborative mechanism effectively balances geometric precision and semantic reasonableness. Furthermore, we introduce a visual-sound decoupling inference paradigm, where the training stage incorporates visual supervisory information, while the inference stage relies solely on acoustic signals for mono-modal localization, thus breaking the traditional dependency on real-time visual data. Experimental results on the AV16.3 dataset demonstrate that ViDAL-Net achieves superior localization performance in visual failure scenarios.