Bridging the Gap: Converting Read Text to Conversational Dialogue
摘要
While the technology of speech processing has gained prominence, it has shed some light on how very difficult it is to convert read speech into a conversational one. It intends to preserve naturalness while at the same time keeping computation requirements within limits enforceable by real-time use. Read speech lacks prosodic variation, which is essential in conversational exchanges. This leads to numerous applications like virtual assistants, customer service, language learning tools, etc. Thus, the current paper proposes a new method, Prosodic Adjustment with Conversational Context (PACC), which seeks to convert read speech into conversational speech in general use. PACC uses advanced deep networks to do adjustments in prosodic features pertaining to intonation, stress, and rhythm. This approach, notably, employs high-fidelity generative adversarial networks (HiFi-GAN) which enhance speech synthesis quality. Our experiments lead to measurable improvement in the anthropic quality and its correctness in setting new standards for conversion tasks on speech. Furthermore, we show that such an approach serves as a means which successfully increases the MOS and can attest to the value of its contribution into other speech conversion scenarios.