错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Voice Separation Using Multi Learning on Squash-Norm Embedding Matrix and Mask

  • Ha Minh Tan,
  • Duc-Quang Vu,
  • Duyen Nguyen Thi,
  • Trang Phung T. Thu

摘要

Recently, deep learning has achieved state-of-the-art performance for many fields e.g., image classification, action recognition, natural language processing, speech recognition, etc. For the speech separation issue, conventional networks directly optimize sources or masks while deep clustering learns embedding matrix training during and clusters an embedding matrix testing during. Deep clustering has obtained some outstanding results in time frequency domain speech separation. Conventional networks have an end-to-end training to take advantage directly signal approximation. The properties and strengths of the traditional networks and deep clustering are combined to improve performances. In this study, we have proposed a multi-learning method by combining a deep clustering and one conventional network. The experiment was conducted on the TIMIT dataset—one of the most standard datasets for speech separation and the results have shown that this method outperforms the component network and state-of-the-art methods recently proposed in terms of common metrics like SDR, SIR, and SAR.