错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adapting Single-Speaker Models for Multi-Speaker Environments

  • M. Vidhyadhar Gowda,
  • M. D. Sujith Kumar,
  • K. Sumukha,
  • Yatin,
  • H. A. Chaya Kumari

摘要

The area of speech, language, and machine learning research has extensively investigated text-to-speech (TTS), sometimes referred to as speech synthesis. Its importance has grown due to its applicability in many different sectors. Neural network-based TTS has significantly improved the quality of synthesized speech, mainly to the breakthroughs in deep learning and artificial intelligence. This study presents an in-depth analysis of neural TTS. We present a text-to-speech (TTS) synthesis system based on neural networks that can produce spoken audio in the voices from multiple speakers, even the ones that are not perceived during training. Our TTS system, with a speaker encoder, Tacotron 2-based synthesis network, and WaveNet-based vocoder, adeptly captures unseen speaker characteristics. It effectively transfers knowledge for synthesizing natural speech from new speakers, emphasizing the importance of diverse training sets and facilitating speech synthesis for dissimilar, novel voices.