<p>The ability to generate dynamic, expressive dance routines that adapt to various musical compositions has broad applications in activity recognition, performance arts, entertainment, virtual reality, and interactive media, offering new avenues for creative professionals and audiences alike. In this article a deep learning framework is developed for music-synchronized dance choreography through modified vision transformers and graph convolutional networks based on Mexican hat wavelet function for position quantization and motion forecasting. More explicitly high-dimensional pose characteristics are extracted from dance video frames using modified vision transformer to generate a skeletal graph, while modified graph convolutional network captures the spatial and temporal relationships between human joints. The process of discretizing continuous pose data is performed by using K-mean clustering and vector quantized variational autoencoders, respectively. The music synchronization beat-aligned loss was optimized, and the best-tuned weight coefficients were found using two variants of the differential evolution algorithm, based on controlled mutation factors <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_21266_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="17" /> </InlineMediaObject> <EquationSource Format="TEX">\(\:\mathcal{F}\)</EquationSource> </InlineEquation> =log-sigmoid () and <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_21266_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="17" /> </InlineMediaObject> <EquationSource Format="TEX">\(\:\mathcal{F}\)</EquationSource> </InlineEquation> =rand(). The proposed architecture with <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_21266_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="17" /> </InlineMediaObject> <EquationSource Format="TEX">\(\:\mathcal{F}\)</EquationSource> </InlineEquation> =log sigmoid () achieves the lowest Fréchet inception distance (FIDk = 32.451, FIDg = 11.219) and music motion correlation of 0.341 demonstrating enhanced motion synthesis in comparison to existed state of art techniques. The mean fitness value of 6.0294 × 10–10 is obtained with an overall classification accuracy of 97.019% in 0.8431G FLOPs for differential evolution algorithm with <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_21266_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="17" /> </InlineMediaObject> <EquationSource Format="TEX">\(\:\mathcal{F}\)</EquationSource> </InlineEquation> log-sigmoid (). The framework may be utilized in AI-generated choreography, virtual dance instruction, and interactive entertainment.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A deep learning based framework for music-synchronized dance choreography with pose quantization and motion prediction for activity recognition

  • Weiwei Fan,
  • Xuerui An

摘要

The ability to generate dynamic, expressive dance routines that adapt to various musical compositions has broad applications in activity recognition, performance arts, entertainment, virtual reality, and interactive media, offering new avenues for creative professionals and audiences alike. In this article a deep learning framework is developed for music-synchronized dance choreography through modified vision transformers and graph convolutional networks based on Mexican hat wavelet function for position quantization and motion forecasting. More explicitly high-dimensional pose characteristics are extracted from dance video frames using modified vision transformer to generate a skeletal graph, while modified graph convolutional network captures the spatial and temporal relationships between human joints. The process of discretizing continuous pose data is performed by using K-mean clustering and vector quantized variational autoencoders, respectively. The music synchronization beat-aligned loss was optimized, and the best-tuned weight coefficients were found using two variants of the differential evolution algorithm, based on controlled mutation factors \(\:\mathcal{F}\) =log-sigmoid () and \(\:\mathcal{F}\) =rand(). The proposed architecture with \(\:\mathcal{F}\) =log sigmoid () achieves the lowest Fréchet inception distance (FIDk = 32.451, FIDg = 11.219) and music motion correlation of 0.341 demonstrating enhanced motion synthesis in comparison to existed state of art techniques. The mean fitness value of 6.0294 × 10–10 is obtained with an overall classification accuracy of 97.019% in 0.8431G FLOPs for differential evolution algorithm with \(\:\mathcal{F}\) log-sigmoid (). The framework may be utilized in AI-generated choreography, virtual dance instruction, and interactive entertainment.