Fusion Attention Graph Convolutional Network with Hyperskeleton for UAV Action Recognition
摘要
Human action recognition (HAR) is a crucial task for inferring human intentions from videos captured by unmanned aerial vehicles (UAVs). However, existing HAR methods lack an attention mechanism capable of capturing motion features across spatial and channel dimensions. In this study, we propose a Spatiotemporal Channel Fusion Attention Module, which is a 3D attention module incorporated after spatial and channel attention. This module collaboratively calculates attention maps across both dimensions, allowing the extraction of highly differentiated neurons and enhancing crucial motion features. To mitigate the impact of occlusions on recognition accuracy, we draw inspiration from how humans infer actions based on unobstructed body parts and propose the Hyperskeleton, a joint cluster spanning multiple body regions that allow for inferring actions based on visible limbs within this cluster. Moreover, leveraging the Gaussian distribution of human motion processes, we introduce a Gaussian Center Enhanced Interpolation Strategy to enrich the frame frequency and information content of videos. Experimental results demonstrate that the proposed model can effectively solving the occlusion problem and achieves state-of-the-art results on the UAV-Human and NTU-RGB+D datasets.