This research focuses on generating nonverbal behavior when an embodied conversational agent (ECA) is the listener. While previous studies on ECAs have focused on generating actions when the agent is the speaker, research on generating actions as a listener has been limited. In particular, it is known that in Japanese dialogue, unlike English dialogue, the correlation between speech and head movements is low. Therefore, a model that considers both linguistic and non-linguistic information is needed. This study constructed a model based on Learning2Listen (L2L) and considered the speaker’s past movements and speaker ID in addition to the speaker’s speech and text information. To improve the model’s performance, VQ-VAE was improved by introducing exponential moving average (EMA) and code reset. In the objective evaluations, VQ-VAE with EMA and code reset showed better performance, and multimodal learning using speech and text was shown to be effective. In the subjective evaluation, the human-likeness of the generated movements and their appropriateness as a listener were evaluated. As a result, the proposed method was shown to be superior to randomly generated movements in terms of appropriateness as a listener. However, compared to actual human movements, the proposed method was suggested to have room for improvement.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generation of Listening Motion of Embodied Conversational Agents Using Speech and Text Information

  • Haruki Ito,
  • Akinori Ito,
  • Takashi Nose

摘要

This research focuses on generating nonverbal behavior when an embodied conversational agent (ECA) is the listener. While previous studies on ECAs have focused on generating actions when the agent is the speaker, research on generating actions as a listener has been limited. In particular, it is known that in Japanese dialogue, unlike English dialogue, the correlation between speech and head movements is low. Therefore, a model that considers both linguistic and non-linguistic information is needed. This study constructed a model based on Learning2Listen (L2L) and considered the speaker’s past movements and speaker ID in addition to the speaker’s speech and text information. To improve the model’s performance, VQ-VAE was improved by introducing exponential moving average (EMA) and code reset. In the objective evaluations, VQ-VAE with EMA and code reset showed better performance, and multimodal learning using speech and text was shown to be effective. In the subjective evaluation, the human-likeness of the generated movements and their appropriateness as a listener were evaluated. As a result, the proposed method was shown to be superior to randomly generated movements in terms of appropriateness as a listener. However, compared to actual human movements, the proposed method was suggested to have room for improvement.