While the technology of speech processing has gained prominence, it has shed some light on how very difficult it is to convert read speech into a conversational one. It intends to preserve naturalness while at the same time keeping computation requirements within limits enforceable by real-time use. Read speech lacks prosodic variation, which is essential in conversational exchanges. This leads to numerous applications like virtual assistants, customer service, language learning tools, etc. Thus, the current paper proposes a new method, Prosodic Adjustment with Conversational Context (PACC), which seeks to convert read speech into conversational speech in general use. PACC uses advanced deep networks to do adjustments in prosodic features pertaining to intonation, stress, and rhythm. This approach, notably, employs high-fidelity generative adversarial networks (HiFi-GAN) which enhance speech synthesis quality. Our experiments lead to measurable improvement in the anthropic quality and its correctness in setting new standards for conversion tasks on speech. Furthermore, we show that such an approach serves as a means which successfully increases the MOS and can attest to the value of its contribution into other speech conversion scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bridging the Gap: Converting Read Text to Conversational Dialogue

  • Parshav Singla,
  • Agnik Banerjee,
  • Aaditya Arora,
  • Shruti Aggarwal,
  • Anil Kumar Verma,
  • C. M. Vikram,
  • Raj Prakash Gohil,
  • Gopal Kumar Agarwal

摘要

While the technology of speech processing has gained prominence, it has shed some light on how very difficult it is to convert read speech into a conversational one. It intends to preserve naturalness while at the same time keeping computation requirements within limits enforceable by real-time use. Read speech lacks prosodic variation, which is essential in conversational exchanges. This leads to numerous applications like virtual assistants, customer service, language learning tools, etc. Thus, the current paper proposes a new method, Prosodic Adjustment with Conversational Context (PACC), which seeks to convert read speech into conversational speech in general use. PACC uses advanced deep networks to do adjustments in prosodic features pertaining to intonation, stress, and rhythm. This approach, notably, employs high-fidelity generative adversarial networks (HiFi-GAN) which enhance speech synthesis quality. Our experiments lead to measurable improvement in the anthropic quality and its correctness in setting new standards for conversion tasks on speech. Furthermore, we show that such an approach serves as a means which successfully increases the MOS and can attest to the value of its contribution into other speech conversion scenarios.