Multi-Scale Co-Attention Network for Fine-Grained Recognition of Video Actions
摘要
To address the three key challenges in fine-grained human action recognition – geometric information loss in keypoint representation, fragmented spatiotemporal features, and cross-modal fusion mismatch – this paper proposes a Multi-Scale Co-Attention network for Fine-Grained Recognition of Video Actions (MSCA-Net). The network pioneers a skeleton keypoint heatmap representation method and employs a dual-branch collaborative architecture to achieve multi-granularity spatiotemporal feature complementarity: the 3D CNN branch captures joint micro-motion patterns, while the 3D Swin Transformer branch models long-range spatiotemporal dependencies. To overcome heterogeneous feature fusion challenges, we design a Multi-scale Co-Attention (MSCA), which enables fine-grained feature enhancement through temporal alignment and dual-dimensional (channel-spatial) gating mechanisms. Experiments on FineGym, NTU60, and NTU120 datasets demonstrate: 1) MSCA-Net achieves state-of-the-art performance in fine-grained action recognition with 95.4% average accuracy, surpassing PoseC3D by 2.2%; 2) The MSCA module realizes the deep fusion of heterogeneous spatio-temporal features based on the feature enhancement module of the triple attention mechanism; 3) Replacing traditional two-stage pose estimators with YOLOv8x-pose achieves real-time inference at 23.4 FPS (RTX3090) while maintaining precision. This study establishes a high-precision, efficient paradigm for fine-grained action understanding in medical rehabilitation assessment and intelligent human-computer interaction scenarios.