<p>3D sound event localization and detection (3D SELD) is a challenging task that integrates sound event detection (SED), direction-of-arrival (DOA) estimation, and sound distance estimation (SDE), aiming to deliver comprehensive spatial information about sound events. To enhance acoustic representation, we propose a Mel-Gammatone Complementary Feature (MGCF), which exploits the complementary properties of Mel and Gammatone filter banks to yield richer and more discriminative spectral features. To generate highly expressive features that can be effectively shared across SED, DOA, and SDE tasks, we design interactive feature aggregation and perception network (IFAP-Net), which aggregates multi-scale spatiotemporal and channel information into a unified high-level representation through an interactive fusion strategy and dual-domain attention refinement. This network comprises an interactive feature aggregation and perception (IFAP) module and a Conformer module, designed to jointly enhance feature aggregation and temporal sequence modeling. Specifically, the IFAP module consists of an interactive feature aggregation block (IFAB) and a time-frequency attention residual block (TFARB). IFAB performs multi-scale fusion across time, frequency, and channel dimensions, while TFARB applies a dual-domain attention mechanism to highlight salient patterns and suppress redundancy. Meanwhile, the Conformer module captures long-range temporal dependencies, further improving sequential modeling. Extensive experiments on the STARSS23 dataset demonstrate that IFAP-Net achieves state-of-the-art (SOTA) performance in feature extraction and aggregation, outperforming existing models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Interactive Feature Aggregation and Perception Network for 3D Sound Event Localization and Detection

  • Yongbo Li,
  • Qinghua Huang

摘要

3D sound event localization and detection (3D SELD) is a challenging task that integrates sound event detection (SED), direction-of-arrival (DOA) estimation, and sound distance estimation (SDE), aiming to deliver comprehensive spatial information about sound events. To enhance acoustic representation, we propose a Mel-Gammatone Complementary Feature (MGCF), which exploits the complementary properties of Mel and Gammatone filter banks to yield richer and more discriminative spectral features. To generate highly expressive features that can be effectively shared across SED, DOA, and SDE tasks, we design interactive feature aggregation and perception network (IFAP-Net), which aggregates multi-scale spatiotemporal and channel information into a unified high-level representation through an interactive fusion strategy and dual-domain attention refinement. This network comprises an interactive feature aggregation and perception (IFAP) module and a Conformer module, designed to jointly enhance feature aggregation and temporal sequence modeling. Specifically, the IFAP module consists of an interactive feature aggregation block (IFAB) and a time-frequency attention residual block (TFARB). IFAB performs multi-scale fusion across time, frequency, and channel dimensions, while TFARB applies a dual-domain attention mechanism to highlight salient patterns and suppress redundancy. Meanwhile, the Conformer module captures long-range temporal dependencies, further improving sequential modeling. Extensive experiments on the STARSS23 dataset demonstrate that IFAP-Net achieves state-of-the-art (SOTA) performance in feature extraction and aggregation, outperforming existing models.