In recent years, video-based Action Quality Assessment (AQA) has gradually become an important research area. However, existing methods tend to treat the video as a whole, ignore the multi-stage dynamic characteristics of the action, and focus mainly on the regression of a single score. This leads to limited prediction accuracy and lack of interpretability of the scoring process. To address these issues, this paper proposes a novel Weight-Aware Scoring and Parsing Network (WASP-Net). The network consists of two key modules: Phase Temporal Parsing Network (PTPN): by modeling the temporal dependencies of video sequences, it achieves precise phase segmentation of actions and generates semantic labels to capture fine-grained action transition features. Weight Aware Scoring Module (WASM): dynamically models the temporal dependencies of action phases and generates phase weights and scores to improve prediction accuracy and interpretability. Extensive experiments on FineDiving and MTL-AQA datasets show that WASP-Net significantly outperforms current state-of-the-art methods in metrics such as Spearman’s rank correlation and relative L2 distance, validating its effectiveness. WASP-Net provides a new perspective for video- based action quality assessment, which not only achieves accurate scoring, but also enhances the scoring interpretability. WASP-Net provides a new perspective for video-based movement quality assessment, which not only realizes accurate scoring but also enhances the interpretability of scoring.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Interpretable Action Quality Assessment with Temporal Parsing

  • Xingquan Cai,
  • Haoyu Song,
  • Yupeng Zhang,
  • Jiatong Li,
  • Haiyan Sun

摘要

In recent years, video-based Action Quality Assessment (AQA) has gradually become an important research area. However, existing methods tend to treat the video as a whole, ignore the multi-stage dynamic characteristics of the action, and focus mainly on the regression of a single score. This leads to limited prediction accuracy and lack of interpretability of the scoring process. To address these issues, this paper proposes a novel Weight-Aware Scoring and Parsing Network (WASP-Net). The network consists of two key modules: Phase Temporal Parsing Network (PTPN): by modeling the temporal dependencies of video sequences, it achieves precise phase segmentation of actions and generates semantic labels to capture fine-grained action transition features. Weight Aware Scoring Module (WASM): dynamically models the temporal dependencies of action phases and generates phase weights and scores to improve prediction accuracy and interpretability. Extensive experiments on FineDiving and MTL-AQA datasets show that WASP-Net significantly outperforms current state-of-the-art methods in metrics such as Spearman’s rank correlation and relative L2 distance, validating its effectiveness. WASP-Net provides a new perspective for video- based action quality assessment, which not only achieves accurate scoring, but also enhances the scoring interpretability. WASP-Net provides a new perspective for video-based movement quality assessment, which not only realizes accurate scoring but also enhances the interpretability of scoring.