Skeleton-Based Action Recognition via Dynamic-Static Partitioning and Cross-Attention Fusion
摘要
The robustness to environmental variations and high processing efficiency have made skeleton-based human action recognition gain significant attention. However, most existing methods overlook the heterogeneity between dynamic and static regions in skeleton sequences and are constrained by traditional feature fusion approaches. To overcome these challenges, we propose two innovative approaches: (1) a gradient-based dynamic-static partitioning mask strategy that combines Grad-CAM and multi-head attention to distinguish and enhance dynamic regions while reducing redundant static information, and (2) a spatio-temporal cross-attention feature fusion strategy that adaptively captures complementary information between dynamic and global features, leading to more discriminative action representations. Extensive experiments on the NTU RGB + D and NTU RGB + D 120 benchmarks demonstrate that our approach significantly improves recognition performance compared to state-of-the-art methods.