In recent years, emotional speech synthesis techniques have attracted considerable interest because of their wide-ranging potential applications. However, when confronted with datasets containing emotional attributes, speech synthesized by traditional methods frequently encounters difficulties, such as a mismatch with the text content or an unnatural expression of emotions. To solve these problems, we developed a straightforward emotional speech synthesis model. This model builds upon the VITS framework and incorporates an emotion prediction module, a prosody prediction module, and a conditional encoder. It can automatically predict emotion labels based on the input text, or manually specify emotions for precise control of emotions. Experimental outcomes indicate improved naturalness and expressiveness in synthesized speech, thus enhancing the overall audio quality.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Expressive Speech Synthesis Enhancement with Conditional Embeddings

  • Fanfan Yan,
  • Maoyu Zhang,
  • Hai Xu,
  • Haoran Ding,
  • Meng Guo

摘要

In recent years, emotional speech synthesis techniques have attracted considerable interest because of their wide-ranging potential applications. However, when confronted with datasets containing emotional attributes, speech synthesized by traditional methods frequently encounters difficulties, such as a mismatch with the text content or an unnatural expression of emotions. To solve these problems, we developed a straightforward emotional speech synthesis model. This model builds upon the VITS framework and incorporates an emotion prediction module, a prosody prediction module, and a conditional encoder. It can automatically predict emotion labels based on the input text, or manually specify emotions for precise control of emotions. Experimental outcomes indicate improved naturalness and expressiveness in synthesized speech, thus enhancing the overall audio quality.