In this work, we propose a neural network approach for speech reconstruction from mel spectrograms, a crucial task in achieving high-quality data after processing speech signals in the time-frequency domain. Specifically, we propose a two-stage deep learning approach based on an overcomplete deep autoencoder (DAE) for the mel filter bank inversion coupled with the deep version of the Griffin-Lim (DeGLI) algorithm for the phase information recovery. After the pre-training of both parts of the architecture, a final fine-tuning on the whole system is performed. Some numerical results, evaluated on the well-known TIMIT dataset, demonstrate the effectiveness of the proposed idea by obtaining a PESQ of 3.996, a STOI equal to 0.994, and a mean opinion score evaluated as 4.15.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Two-Stage Neural Network for Speech Signal Reconstruction from Mel Spectrograms

  • Filippo Villani,
  • Michele Scarpiniti,
  • Aurelio Uncini

摘要

In this work, we propose a neural network approach for speech reconstruction from mel spectrograms, a crucial task in achieving high-quality data after processing speech signals in the time-frequency domain. Specifically, we propose a two-stage deep learning approach based on an overcomplete deep autoencoder (DAE) for the mel filter bank inversion coupled with the deep version of the Griffin-Lim (DeGLI) algorithm for the phase information recovery. After the pre-training of both parts of the architecture, a final fine-tuning on the whole system is performed. Some numerical results, evaluated on the well-known TIMIT dataset, demonstrate the effectiveness of the proposed idea by obtaining a PESQ of 3.996, a STOI equal to 0.994, and a mean opinion score evaluated as 4.15.