错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PiCo-VITS: Leveraging Pitch Contours for Fine-Grained Emotional Speech Synthesis

  • Kwan-yeung Wong,
  • Fu-lai Chung

摘要

Text-to-speech (TTS) research has made significant progress in achieving human-like interpretability. Yet, a noticeable research gap persists regarding emotional expressiveness in synthesized speech. While some existing studies address this through techniques such as voice conversion and textual context modeling, precise control over emotion composition in speech synthesis remains challenging. Moreover, those techniques typically operate at sentence-level, precluding fine-grained control over emotion transitions. Emotion composition in utterances is multifaceted, dictated by a multitude of linguistic and acoustic features; and among these features, pitch plays a crucial role, with distinct pitch contours often signaling specific emotions at different points in a spoken sentence. Aiming for more controllable emotional speech synthesis, we propose PiCo-VITS, an end-to-end TTS model architecture that leverages pitch contours in conjunction with latent features. Experimental results demonstrate the efficacy of the proposed model in synthesizing speech that conveys mixed emotions. Notably, the model allows both the desired emotions and the emotion transition patterns to be specified, while maintaining intelligibility comparable to state-of-the-art techniques.