错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Legato: Articulation and Dynamic-Aware Dance Generation via VQ-VAE

  • Sun Young Park,
  • Ji Su Park,
  • Jin Gon Shon

摘要

Music-driven dance generation is a popular research topic in computer vision with applications in choreography creation and performance production. While existing research has primarily focused on rhythmic features such as beats and tempo, expressive elements like dynamics (loudness variations) and articulation (note connectivity) have not been adequately addressed. This work proposes Legato, a VQ-VAE architecture that explicitly separates musical expressive elements from pose information. We extract musical features including RMS Energy for dynamics, melody-based Key Overlap Ratio for articulation, and Beat Mask for rhythm synchronization. Among these, dynamics and articulation are quantized into separate codebooks alongside pose codebooks (upper/lower body), forming a total of four independent codebooks, while beat information is directly fed into Motion GPT without a separate codebook to handle rhythm synchronization. We also introduce Adaptive Jerk Loss, which adaptively controls motion smoothness according to musical context. To address musical expressiveness in dance generation, we provide explicit modeling of dynamics and articulation.