Dance and music, as universal languages of emotion and expression, have been integral practices throughout human history, often associated with social and religious ceremonies. Manually animating a dancing person based on music is a challenging task that requires skill, time, and effort. However, with the help of an artificial intelligence (AI) model, dances can be generated automatically in response to music. Despite significant advancements in motion generation using AI techniques such as transformers, diffusion models, and GANs, challenges remain because existing frameworks primarily aim to produce movements that appear plausible in a general sense, rather than fully realistic. We developed our model with a specific goal: to understand the rhythmic link between music and motion and build specific components to learn the relationship between these features and then utilize that in the generative model. To achieve this, we employ a state-of-the-art diffusion-style model to create dance sequences. We then introduce two sub-models: the Fusion Sync Classifier and Fusion Sync Enhancer. These sub-models, when integrated into the main model, Rhythm Fusion, ensure audio-video synchronization and facilitate the alignment and correlation between motion and music. Through the use of quantitative metrics, we show that our model outperforms other state-of-the-art models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Rhythm Fusion: Synchronizing Audio and Motion Features for Music-Driven Dance Generation

  • Nuha Aldausari,
  • Gelareh Mohammadi,
  • David Cooper

摘要

Dance and music, as universal languages of emotion and expression, have been integral practices throughout human history, often associated with social and religious ceremonies. Manually animating a dancing person based on music is a challenging task that requires skill, time, and effort. However, with the help of an artificial intelligence (AI) model, dances can be generated automatically in response to music. Despite significant advancements in motion generation using AI techniques such as transformers, diffusion models, and GANs, challenges remain because existing frameworks primarily aim to produce movements that appear plausible in a general sense, rather than fully realistic. We developed our model with a specific goal: to understand the rhythmic link between music and motion and build specific components to learn the relationship between these features and then utilize that in the generative model. To achieve this, we employ a state-of-the-art diffusion-style model to create dance sequences. We then introduce two sub-models: the Fusion Sync Classifier and Fusion Sync Enhancer. These sub-models, when integrated into the main model, Rhythm Fusion, ensure audio-video synchronization and facilitate the alignment and correlation between motion and music. Through the use of quantitative metrics, we show that our model outperforms other state-of-the-art models.