The robustness to environmental variations and high processing efficiency have made skeleton-based human action recognition gain significant attention. However, most existing methods overlook the heterogeneity between dynamic and static regions in skeleton sequences and are constrained by traditional feature fusion approaches. To overcome these challenges, we propose two innovative approaches: (1) a gradient-based dynamic-static partitioning mask strategy that combines Grad-CAM and multi-head attention to distinguish and enhance dynamic regions while reducing redundant static information, and (2) a spatio-temporal cross-attention feature fusion strategy that adaptively captures complementary information between dynamic and global features, leading to more discriminative action representations. Extensive experiments on the NTU RGB + D and NTU RGB + D 120 benchmarks demonstrate that our approach significantly improves recognition performance compared to state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Skeleton-Based Action Recognition via Dynamic-Static Partitioning and Cross-Attention Fusion

  • Hui Cao,
  • Yuanyuan Wang,
  • Tingwei Wang

摘要

The robustness to environmental variations and high processing efficiency have made skeleton-based human action recognition gain significant attention. However, most existing methods overlook the heterogeneity between dynamic and static regions in skeleton sequences and are constrained by traditional feature fusion approaches. To overcome these challenges, we propose two innovative approaches: (1) a gradient-based dynamic-static partitioning mask strategy that combines Grad-CAM and multi-head attention to distinguish and enhance dynamic regions while reducing redundant static information, and (2) a spatio-temporal cross-attention feature fusion strategy that adaptively captures complementary information between dynamic and global features, leading to more discriminative action representations. Extensive experiments on the NTU RGB + D and NTU RGB + D 120 benchmarks demonstrate that our approach significantly improves recognition performance compared to state-of-the-art methods.