Multi-channel speech enhancement aims to extract the clean speech signal from noisy mixture recordings captured by the microphone array. In this work, we propose a new end-to-end framework – time-frequency graph transformer network (TFGT-Net), which views the embedded space of the encoder network as a potential graphical representation. The TFGT-Net combines the advantages of graph neural network (GNN) and transformer, where GNN is used to learn the spatial correlations between different channels in the dimensions of time and frequency, and the output of the graph convolutional layer is fed to a transformer block to further capture the time-frequency long-term dependence to estimate the complex spectral mask. Then the obtained mask is multiplied with the target microphone spectrogram to get the enhanced speech. Moreover, we perform experiments of multi-channel speech enhancement based on the DNS-Challenge dataset, and the results show that our proposed framework outperforms several recent neural enhancement models in multiple evaluation indicators.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TFGT-Net: Time-Frequency Graph Transformer Network For Multi-channel Speech Enhancement

  • Yongshuai Li,
  • Songlin Sun,
  • Chenwei Wang

摘要

Multi-channel speech enhancement aims to extract the clean speech signal from noisy mixture recordings captured by the microphone array. In this work, we propose a new end-to-end framework – time-frequency graph transformer network (TFGT-Net), which views the embedded space of the encoder network as a potential graphical representation. The TFGT-Net combines the advantages of graph neural network (GNN) and transformer, where GNN is used to learn the spatial correlations between different channels in the dimensions of time and frequency, and the output of the graph convolutional layer is fed to a transformer block to further capture the time-frequency long-term dependence to estimate the complex spectral mask. Then the obtained mask is multiplied with the target microphone spectrogram to get the enhanced speech. Moreover, we perform experiments of multi-channel speech enhancement based on the DNS-Challenge dataset, and the results show that our proposed framework outperforms several recent neural enhancement models in multiple evaluation indicators.