PiCo-VITS: Leveraging Pitch Contours for Fine-Grained Emotional Speech Synthesis
摘要
Text-to-speech (TTS) research has made significant progress in achieving human-like interpretability. Yet, a noticeable research gap persists regarding emotional expressiveness in synthesized speech. While some existing studies address this through techniques such as voice conversion and textual context modeling, precise control over emotion composition in speech synthesis remains challenging. Moreover, those techniques typically operate at sentence-level, precluding fine-grained control over emotion transitions. Emotion composition in utterances is multifaceted, dictated by a multitude of linguistic and acoustic features; and among these features, pitch plays a crucial role, with distinct pitch contours often signaling specific emotions at different points in a spoken sentence. Aiming for more controllable emotional speech synthesis, we propose PiCo-VITS, an end-to-end TTS model architecture that leverages pitch contours in conjunction with latent features. Experimental results demonstrate the efficacy of the proposed model in synthesizing speech that conveys mixed emotions. Notably, the model allows both the desired emotions and the emotion transition patterns to be specified, while maintaining intelligibility comparable to state-of-the-art techniques.