Conversational gesture generation holds pivotal importance in enhancing the natural flow of digital human interaction. However, the inherent weak correlation between gestures and language poses a significant challenge for automated gesture generation systems. Existing methods have shortcomings in the quality of generated gestures and inference speed. In this paper, we propose an optimized method for conversational gesture generation. By training a more powerful motion feature extraction network, the connotation of different motions can be accurately grasped. Moreover, we utilize a two-level cascaded generator, which initially produces broad movements such as arm gestures, followed by generating finer gestures like finger movements. To enhance network convergence and stability, we integrate multiple loss terms and implement refined training strategies. These measures collectively contribute to the improved quality of the final generated gestures. Additionally, we employ GRU as the backbone network, striking a balance between generation quality and computational efficiency. Both subjective and objective experiments demonstrate that our method achieves a remarkable inference speed while generating high-quality gestures.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimized Conversational Gesture Generation with Enhanced Motion Feature Extraction and Cascaded Generator

  • Xiang Wang,
  • Yifeng Peng,
  • Zhaoxiang Liu,
  • Shijie Dong,
  • Ruitao Liu,
  • Kai Wang,
  • Shiguo Lian

摘要

Conversational gesture generation holds pivotal importance in enhancing the natural flow of digital human interaction. However, the inherent weak correlation between gestures and language poses a significant challenge for automated gesture generation systems. Existing methods have shortcomings in the quality of generated gestures and inference speed. In this paper, we propose an optimized method for conversational gesture generation. By training a more powerful motion feature extraction network, the connotation of different motions can be accurately grasped. Moreover, we utilize a two-level cascaded generator, which initially produces broad movements such as arm gestures, followed by generating finer gestures like finger movements. To enhance network convergence and stability, we integrate multiple loss terms and implement refined training strategies. These measures collectively contribute to the improved quality of the final generated gestures. Additionally, we employ GRU as the backbone network, striking a balance between generation quality and computational efficiency. Both subjective and objective experiments demonstrate that our method achieves a remarkable inference speed while generating high-quality gestures.