错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fine-Grained Style Control in VITS-Based Text-to-Speech Synthesis

  • Zhong Huihang,
  • Dengfeng Ke,
  • Li Ya,
  • Wenhan Yao,
  • Wenqian Bao

摘要

In this paper, a fine-grained style controllable speech synthesis model based on VITS is presented. To achieve fine-grained emotional speech, global and local emotion features are extracted using GST and LST, respectively. A multi head cross-attention mechanism is used to align the text with emotion features, achieving prosodic control on the text side. Experiments demonstrate that the original text encoder of VITS is not capable of handling both text and emotion features simultaneously. Therefore, an emotion-text encoder is proposed, which significantly improves the ability of the prior encoder and enables fine-grained emotion speech synthesis. Results show that the system achieves the highest levels of naturalness, style conversion ability, and speech quality in synthesized speech. Audio samples are publicly available ( https://lunar333.github.io ).