Explorative Study on the Non-verbal Backchannel Prediction Model for Human-Robot Interaction
摘要
Previous studies on backchannel prediction model have suggested that replicating human backchannel can enhance user’s human-robot interaction experience. In this study, we propose a real-time non-verbal backchannel prediction model which utilizes both an acoustic feature and a temporal feature. Our goal is to improve the quality of robot’s backchannel and user’s experience. To conduct this research, we collected a human-human interview dataset. Using this dataset, we proceeded to develop three distinct backchannel prediction models: a temporal, an acoustic, and a mixed (temporal & acoustic) model. Subsequently, we conducted a user study to compare the perception of robot implemented with the three models. The results demonstarted that the robot employing the mixed model was preferred by participants and exhibited moderate frequency of backchannel. These results emphasize the advantages of incorporating acoustic and temporal features in developing backchannel prediction model to enhance the quality of human-robot interactions, specifically with regards to backchannel frequency and timing.