In conversational Text-to-Speech task, we aim to synthesize speech with correct linguistic and emotional rhythm in the conversational context. However, in previous studies, the impact of fine-grained contextual historical information on emotional communication during dialogue has not been given enough consideration. This may cause the loss of subtle emotional clues. To address this issue, in this paper, we propose an approach to predict emotions using fine-grained historical context. Specifically, we design a heterogeneous graph to obtain fine-grained emotional clues from the context of a conversation. Further, in order to extract more accurate and delicate emotional features, we introduce a simple and effective emotion extractor. We integrate the emotional features into acoustic features to synthesize speech. It contains rich emotional expressions in the synthesized current discourse that match the emotional expressions in the context. We conduct experiments on the existing conversational datasets DailyTalk and DailyDialogue. The results show that the method proposed in this paper has excellent conversational speech synthesis quality.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predicting Emotions in Conversational Speech Synthesis Using Fine-Grained Historical Context

  • Hanwei Li,
  • Qiulin Li,
  • Xin Dong,
  • Qun Yang

摘要

In conversational Text-to-Speech task, we aim to synthesize speech with correct linguistic and emotional rhythm in the conversational context. However, in previous studies, the impact of fine-grained contextual historical information on emotional communication during dialogue has not been given enough consideration. This may cause the loss of subtle emotional clues. To address this issue, in this paper, we propose an approach to predict emotions using fine-grained historical context. Specifically, we design a heterogeneous graph to obtain fine-grained emotional clues from the context of a conversation. Further, in order to extract more accurate and delicate emotional features, we introduce a simple and effective emotion extractor. We integrate the emotional features into acoustic features to synthesize speech. It contains rich emotional expressions in the synthesized current discourse that match the emotional expressions in the context. We conduct experiments on the existing conversational datasets DailyTalk and DailyDialogue. The results show that the method proposed in this paper has excellent conversational speech synthesis quality.