Tacotron2 [1] is a speech synthesis model that can convert text into Mel Spectrograms [1] and generate natural-sounding speech waveforms, sound close to the human voice. We hope that this technology can provide a valuable resource for those who struggle with public speaking or are temporarily unable to speak. This study uses WaveGlow [2] to convert Mel Spectrograms [1] into audible sound waves, read the Japanese text, and output the speech with Japanese voice, actor Saori Hayami as the voice actor. We chose to use real human voice actors to train the speech system to improve the naturalness and clarity of Japanese synthesized speech, and to make a clear distinction from AI voices such as Google Assistant and Apple Siri.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Text to Human-Like Speech Using Tacotron-Based TTS Model

  • Wei-Chen Wu,
  • Zhi-Xun Cai,
  • Chin-Feng Lee

摘要

Tacotron2 [1] is a speech synthesis model that can convert text into Mel Spectrograms [1] and generate natural-sounding speech waveforms, sound close to the human voice. We hope that this technology can provide a valuable resource for those who struggle with public speaking or are temporarily unable to speak. This study uses WaveGlow [2] to convert Mel Spectrograms [1] into audible sound waves, read the Japanese text, and output the speech with Japanese voice, actor Saori Hayami as the voice actor. We chose to use real human voice actors to train the speech system to improve the naturalness and clarity of Japanese synthesized speech, and to make a clear distinction from AI voices such as Google Assistant and Apple Siri.