Comparative Analysis on Speech Driven Gesture Generation
摘要
This study presents a comparative analysis of various deep- learning methods for gesture generation from speech, contributing to the field of human-agent interaction. The primary focus lies in advancing interactions with virtual agents and robots through the application of varied representation learning approaches. The methodologies employ Gated Recurrent Unit (GRU), Bidirectional-GRU, and Multi-head attention techniques to build a concise representation of human gestures, acting as motion encoders and decoders. Incorporating a pre-existing gesture generation approach, the established groundwork through the utilization of a network named Speech E is leveraged. Furthermore, an in-depth analysis of diverse speech feature inputs’ impact on the model’s performance is pursued. This network is designed to efficiently transform speech input into a gesture representation with reduced dimensions. The influence of different speech feature inputs on the model’s performance is explored. The integration of GRU, Bidirectional-GRU, and Multi-Head Attention methods allow a thorough evaluation of their effectiveness in translating speech into corresponding gestures. This study provides insights into the potential of these techniques, and their implications for creating more intuitive and responsive virtual agents and robots. Notably, our investigations exclusively utilize Mel-Frequency Cepstral Coefficients (MFCCs) as features, revealing the optimal performance of our models. This careful feature selection shines brightly in our results, where our Bi-directional GRU model outshines the rest.