Lightweight Multispectral Skeleton and Multi-stream Graph Attention Networks for Enhanced Action Prediction with Multiple Modalities
摘要
Human action recognition methods often focus on extracting structural and temporal information from skeleton-based graphs. However, these approaches struggle with effectively capturing and processing extensive information during action transitions. To overcome this limitation, we propose LMS-GAT, a novel approach that facilitates information exchange through node concentration and diffusion across structural and temporal dimensions. By selectively suppressing and reinstating the representations of structural nodes for each specific action, and utilizing hierarchical shifted temporal windows for assessing temporal information, LMS-GAT addresses the challenge of dynamic changes in action recognition. Experimental evaluation on NTU RGB+D 60 and 120 datasets shows that LMS-GAT outperforms state-of-the-art methods in terms of prediction accuracy. This highlights the efficacy of our approach in capturing and recognizing human actions with improved performance.