Abstract <p>Generating context-appropriate human motions given an audio input is hard but has a wide range of application uses in computer games, virtual reality systems and in the entertainment industry. There are many challenges for audio-based motion synthesis: (1) human motions are highly diverse and expressive which can carry different motion content with various styles in both time and space; and (2) human motions need to be well coordinated and synchronized with other modalities to appear consistent and convincing. In this paper, we regard human motions as a form of body language expressions and propose a two-level hierarchical framework as a solution to audio-driven motion synthesis tasks. During preprocessing, long motion sequences in the dataset are segmented into motion words as building blocks for motion composition, with word embeddings extracted through a pre-trained autoencoder. Our high-level planner is a transformer-based sequence model trained to predict the next word embedding given previous embeddings and optional accompanying audio as input. The low-level performance implementer takes the predicted word embeddings and the accompanying audio, identifies the motion style that best matches the audio influence, and synchronizes the motion to the audio beats. We demonstrate the successful application of our framework to two typical audio-based motion synthesis tasks: music-driven dance choreography and prosody-driven gesture generation. For both applications, results show that our framework generates high quality, diverse motions that are well synchronized to audio. We conclude by discussing future work to further enhance audio-driven motion synthesis research.</p> Graphic Abstract <p></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Audio2Moves: Two-Level Hierarchical Framework for Audio-Driven Human Motion Synthesis

  • Yanbo Cheng,
  • Nada Elmasry,
  • Yingying Wang

摘要

Abstract

Generating context-appropriate human motions given an audio input is hard but has a wide range of application uses in computer games, virtual reality systems and in the entertainment industry. There are many challenges for audio-based motion synthesis: (1) human motions are highly diverse and expressive which can carry different motion content with various styles in both time and space; and (2) human motions need to be well coordinated and synchronized with other modalities to appear consistent and convincing. In this paper, we regard human motions as a form of body language expressions and propose a two-level hierarchical framework as a solution to audio-driven motion synthesis tasks. During preprocessing, long motion sequences in the dataset are segmented into motion words as building blocks for motion composition, with word embeddings extracted through a pre-trained autoencoder. Our high-level planner is a transformer-based sequence model trained to predict the next word embedding given previous embeddings and optional accompanying audio as input. The low-level performance implementer takes the predicted word embeddings and the accompanying audio, identifies the motion style that best matches the audio influence, and synchronizes the motion to the audio beats. We demonstrate the successful application of our framework to two typical audio-based motion synthesis tasks: music-driven dance choreography and prosody-driven gesture generation. For both applications, results show that our framework generates high quality, diverse motions that are well synchronized to audio. We conclude by discussing future work to further enhance audio-driven motion synthesis research.

Graphic Abstract