An Interactive Feature Aggregation and Perception Network for 3D Sound Event Localization and Detection
摘要
3D sound event localization and detection (3D SELD) is a challenging task that integrates sound event detection (SED), direction-of-arrival (DOA) estimation, and sound distance estimation (SDE), aiming to deliver comprehensive spatial information about sound events. To enhance acoustic representation, we propose a Mel-Gammatone Complementary Feature (MGCF), which exploits the complementary properties of Mel and Gammatone filter banks to yield richer and more discriminative spectral features. To generate highly expressive features that can be effectively shared across SED, DOA, and SDE tasks, we design interactive feature aggregation and perception network (IFAP-Net), which aggregates multi-scale spatiotemporal and channel information into a unified high-level representation through an interactive fusion strategy and dual-domain attention refinement. This network comprises an interactive feature aggregation and perception (IFAP) module and a Conformer module, designed to jointly enhance feature aggregation and temporal sequence modeling. Specifically, the IFAP module consists of an interactive feature aggregation block (IFAB) and a time-frequency attention residual block (TFARB). IFAB performs multi-scale fusion across time, frequency, and channel dimensions, while TFARB applies a dual-domain attention mechanism to highlight salient patterns and suppress redundancy. Meanwhile, the Conformer module captures long-range temporal dependencies, further improving sequential modeling. Extensive experiments on the STARSS23 dataset demonstrate that IFAP-Net achieves state-of-the-art (SOTA) performance in feature extraction and aggregation, outperforming existing models.