The purpose of speech separation algorithms is to solve the “cocktail party problem” by separating the voice of the main speaker from the mixed speech. Currently, most speech separation algorithms based on deep learning usually require permutation invariant training due to the need to separate multiple target sources of speech at once. However, this paper proposes a new multi-channel target direction speech separation model, which only separates speech in the target direction without separating all sources together, thereby directly avoiding permutation invariant training. In this model, multi-channel spatial features guide the model to learn speech related information in the target direction. And in the learning of temporal information, unlike other speech separation models, this paper uses an efficient convolutional network, U-shaped convolutional block network, as the separator of the model. This paper proves that this temporal model is more efficient in learning spatiotemporal features. Finally, the model is tested using the multi-channel speech dataset created. The experiment shows that the proposed model has better separation performance compared to other single channel speech separation models and multi-channel speech separation models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Multi-channel Target Direction Speech Separation Model Based on Efficient Convolutional Networks

  • Zaixin Chen,
  • Yang Mao,
  • Tihua Yan,
  • Biqing Li

摘要

The purpose of speech separation algorithms is to solve the “cocktail party problem” by separating the voice of the main speaker from the mixed speech. Currently, most speech separation algorithms based on deep learning usually require permutation invariant training due to the need to separate multiple target sources of speech at once. However, this paper proposes a new multi-channel target direction speech separation model, which only separates speech in the target direction without separating all sources together, thereby directly avoiding permutation invariant training. In this model, multi-channel spatial features guide the model to learn speech related information in the target direction. And in the learning of temporal information, unlike other speech separation models, this paper uses an efficient convolutional network, U-shaped convolutional block network, as the separator of the model. This paper proves that this temporal model is more efficient in learning spatiotemporal features. Finally, the model is tested using the multi-channel speech dataset created. The experiment shows that the proposed model has better separation performance compared to other single channel speech separation models and multi-channel speech separation models.