Decoupled pre-training and multi-modality fusion for fine-grained action quality assessment
摘要
Action Quality Assessment (AQA) for diving demands integrated features of human movements and environmental interactions, yet existing works face challenges of insufficient skeleton data and underutilized multi-modal complementarity. To address these, this study first proposes the FineDiving-Pose dataset. By transforming skeletons into stacked pseudo-heatmaps, we achieve unified spatial-temporal modeling with RGB data, avoiding reliance on specialized graph neural networks. We then propose a staged pre-training strategy (decoupling action recognition and AQA) and optimize temporal sampling (replacing redundant two-step regular sampling), enhancing baseline TSA performance while reducing input length. Additionally, we modify the Non-local operator into a dual-input attention fusion block for early RGB-Pose interaction. Qualitative analysis via spatiotemporal attention heatmaps (Dive 405B) shows Pose excels at interference-free human tracking, while RGB captures environmental cues (e.g., water splashes). Quantitative experiments on FineDiving-Pose demonstrate the method outperforms single-modal/baseline models in