Insect noise (e.g., cicadas) is the primary noise interference in outdoor audio collection during summer and autumn. Current speech enhancement models suffer from ineffective multi-scale feature utilization and redundant information aggregation, hindering speech recovery. In order to more efficiently solve the interference of insect noise, this paper introduces a Multi-scale Adaptive Feature Sparse Network (MAFSNet). The Multi-scale Adaptive Sparse Conformer (MASC), serving as the cornerstone design of MAFSNet, comprises three pivotal constituent modules. Specifically, Multi-scale Adaptive Sparse Attention (MASA) module differentially processes multi-scale features at different levels, and then uses a dual-branch self-attention for adaptive redundancy reduction. Channel-guided Branch Modulation (CBM) module combines channel features with gating mechanisms to suppress irrelevant features of different branches. Meanwhile, Dual Scale Enhancement Fusion (DSEF) module implements learnable weighting for optimized feature fusion. The quantitative evaluations of the Insect Noise dataset and the Voice Bank+DEMAND dataset demonstrate that MAFSNet outperformed other models and achieved excellent results in objective evaluation metrics.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-scale Adaptive Feature Sparse Network for Speech Enhancement in Insect Noise Environments

  • Chao Zhang,
  • Jian He,
  • Xiangping Gao,
  • Dongmei Zheng

摘要

Insect noise (e.g., cicadas) is the primary noise interference in outdoor audio collection during summer and autumn. Current speech enhancement models suffer from ineffective multi-scale feature utilization and redundant information aggregation, hindering speech recovery. In order to more efficiently solve the interference of insect noise, this paper introduces a Multi-scale Adaptive Feature Sparse Network (MAFSNet). The Multi-scale Adaptive Sparse Conformer (MASC), serving as the cornerstone design of MAFSNet, comprises three pivotal constituent modules. Specifically, Multi-scale Adaptive Sparse Attention (MASA) module differentially processes multi-scale features at different levels, and then uses a dual-branch self-attention for adaptive redundancy reduction. Channel-guided Branch Modulation (CBM) module combines channel features with gating mechanisms to suppress irrelevant features of different branches. Meanwhile, Dual Scale Enhancement Fusion (DSEF) module implements learnable weighting for optimized feature fusion. The quantitative evaluations of the Insect Noise dataset and the Voice Bank+DEMAND dataset demonstrate that MAFSNet outperformed other models and achieved excellent results in objective evaluation metrics.