Recently, a state space model (SSM) called Mamba achieves remarkable performance, particularly in vision tasks, due to its faster inference and shows significant potential in handling tasks involving long sequences and autoregressive properties. To explore the efficient Mamba-Net for speech enhancement and alleviate the complexity of the system, we proposed a Mamba-based CMGAN (M-CMGAN) for speech enhancement in the time-frequency domain. In this work, we combined Mamba with a cepstrum frequency block (CFB) and the CMGAN backbone to efficiently save GPU memory and complement spatial information. Three functional blocks (FFB, MMB, and GMB) were introduced to facilitate information exchange. The experiments on the Voice Bank+DEMAND dataset prove that the proposed method reduces one-third of the training time and possibly more. Compared to the reproduced CMGAN, our model improves about 7 \(\%\) more SSNR and 15.4 \(\%\) less required estimated size. We additionally found that MMB used with only one layer of TS-Conformer can release one-third more parameters, it spends less than half the time before but the PESQ only decreased by 0.05. By removing FFB, we obtain around 10 \(\%\) improvement for SSNR and reduced 31.9 \(\%\) total estimation, which is equivalent to saving 39.7 \(\%\) training time, with the PESQ reaching 3.42, outperforming the baseline.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

M-CMGAN: Attempting to Use Mamba on Speech Enhancement

  • Yujie Xiong,
  • Zhihua Huang

摘要

Recently, a state space model (SSM) called Mamba achieves remarkable performance, particularly in vision tasks, due to its faster inference and shows significant potential in handling tasks involving long sequences and autoregressive properties. To explore the efficient Mamba-Net for speech enhancement and alleviate the complexity of the system, we proposed a Mamba-based CMGAN (M-CMGAN) for speech enhancement in the time-frequency domain. In this work, we combined Mamba with a cepstrum frequency block (CFB) and the CMGAN backbone to efficiently save GPU memory and complement spatial information. Three functional blocks (FFB, MMB, and GMB) were introduced to facilitate information exchange. The experiments on the Voice Bank+DEMAND dataset prove that the proposed method reduces one-third of the training time and possibly more. Compared to the reproduced CMGAN, our model improves about 7 \(\%\) more SSNR and 15.4 \(\%\) less required estimated size. We additionally found that MMB used with only one layer of TS-Conformer can release one-third more parameters, it spends less than half the time before but the PESQ only decreased by 0.05. By removing FFB, we obtain around 10 \(\%\) improvement for SSNR and reduced 31.9 \(\%\) total estimation, which is equivalent to saving 39.7 \(\%\) training time, with the PESQ reaching 3.42, outperforming the baseline.