<p>Action Quality Assessment (AQA) for diving demands integrated features of human movements and environmental interactions, yet existing works face challenges of insufficient skeleton data and underutilized multi-modal complementarity. To address these, this study first proposes the FineDiving-Pose dataset. By transforming skeletons into stacked pseudo-heatmaps, we achieve unified spatial-temporal modeling with RGB data, avoiding reliance on specialized graph neural networks. We then propose a staged pre-training strategy (decoupling action recognition and AQA) and optimize temporal sampling (replacing redundant two-step regular sampling), enhancing baseline TSA performance while reducing input length. Additionally, we modify the Non-local operator into a dual-input attention fusion block for early RGB-Pose interaction. Qualitative analysis via spatiotemporal attention heatmaps (Dive 405B) shows Pose excels at interference-free human tracking, while RGB captures environmental cues (e.g., water splashes). Quantitative experiments on FineDiving-Pose demonstrate the method outperforms single-modal/baseline models in <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\varvec{\rho }\)</EquationSource> </InlineEquation> and <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\varvec{R}\)</EquationSource> </InlineEquation>-<InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\varvec{l}_{\varvec{2}}\)</EquationSource> </InlineEquation>, with Pose’s sparse pseudo-heatmaps ensuring faster inference than RGB. Limitations and future directions (e.g., feature alignment for pre-trained weight reuse, generalization to other AQA domains) are also discussed, providing a lightweight, efficient framework for multi-modal AQA.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Decoupled pre-training and multi-modality fusion for fine-grained action quality assessment

  • Jiahao Guan,
  • Chengming Liu,
  • Lei Shi

摘要

Action Quality Assessment (AQA) for diving demands integrated features of human movements and environmental interactions, yet existing works face challenges of insufficient skeleton data and underutilized multi-modal complementarity. To address these, this study first proposes the FineDiving-Pose dataset. By transforming skeletons into stacked pseudo-heatmaps, we achieve unified spatial-temporal modeling with RGB data, avoiding reliance on specialized graph neural networks. We then propose a staged pre-training strategy (decoupling action recognition and AQA) and optimize temporal sampling (replacing redundant two-step regular sampling), enhancing baseline TSA performance while reducing input length. Additionally, we modify the Non-local operator into a dual-input attention fusion block for early RGB-Pose interaction. Qualitative analysis via spatiotemporal attention heatmaps (Dive 405B) shows Pose excels at interference-free human tracking, while RGB captures environmental cues (e.g., water splashes). Quantitative experiments on FineDiving-Pose demonstrate the method outperforms single-modal/baseline models in \(\varvec{\rho }\) and \(\varvec{R}\) - \(\varvec{l}_{\varvec{2}}\) , with Pose’s sparse pseudo-heatmaps ensuring faster inference than RGB. Limitations and future directions (e.g., feature alignment for pre-trained weight reuse, generalization to other AQA domains) are also discussed, providing a lightweight, efficient framework for multi-modal AQA.