A Text to Human-Like Speech Using Tacotron-Based TTS Model
摘要
Tacotron2 [1] is a speech synthesis model that can convert text into Mel Spectrograms [1] and generate natural-sounding speech waveforms, sound close to the human voice. We hope that this technology can provide a valuable resource for those who struggle with public speaking or are temporarily unable to speak. This study uses WaveGlow [2] to convert Mel Spectrograms [1] into audible sound waves, read the Japanese text, and output the speech with Japanese voice, actor Saori Hayami as the voice actor. We chose to use real human voice actors to train the speech system to improve the naturalness and clarity of Japanese synthesized speech, and to make a clear distinction from AI voices such as Google Assistant and Apple Siri.