Multi-modal Variable-Channel Spatial-Temporal Semantic Action Recognition Network
摘要
Multimodal action recognition, especially the fusion of image and skeleton data, has emerged as the prevailing approach in the field of action recognition. However, existing models often lack the spatial-temporal discriminative ability required for fine-grained recognition tasks. To address this issue, we introduce a flexible attention block called Variable-Channel Spatial-Temporal attention (VCSTA) to enhance the discriminative capacity of spatial-temporal connections. Based on VCSTA, we propose a novel multimodal variable-channel Spatial-Temporal semantic action recognition network (MMARN). MMARN utilizes a combination of spatial-temporal embedding loss and global loss to improve the model’s understanding of action changes and semantic information in videos, resulting in more precise prediction and analysis. The experimental results show that our multimodal variable-channel spatial-temporal semantic action recognition network achieves 98.3% and 89.9% accuracy in classifying actions on the large-scale human activity datasets NTU-RGB+D 60 and NTU-RGB+D 120 respectively.