Generation of Listener’s Facial Response Using Cross-Modal Mapping of Speaker’s Expression
摘要
In human communication, we use non-verbal cues such as facial expressions to convey intentions and emotional states.These non-verbal elements play important roles in enhancing the depth and understanding of conversations. In a dyadic interaction, there are roles of a speaker and a listener, which vary depending on the phase of the interaction. Listeners often employ variations in facial expressions and nods as a means to signal their engagement and response to the speaker.There has been research on the task of developing machine learning models that generate listener facial responses from the speaker’s facial expressions and speech information in recent years. However, although these models provide a framework for handling both facial images and audio information as a speaker’s expressions, they do not focus much on the relationship between the speaker’s facial expressions and voice expressions. Since facial and voice expressions are often related when conducting a conversation, a model that takes these into account could be useful.Therefore, we focused on the methods used in the cross-modal generalization task. These are methods for learning a unified discrete representation from a combination of multimodal data by mapping them into the same representation space. In this study, we examine the effectiveness of combining the conventionally used LSTM-based model with a unified mapping of multimodal data on the task of generating the listener’s facial expressions from the speaker’s facial expression and voice during communication.