错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-modal Variable-Channel Spatial-Temporal Semantic Action Recognition Network

  • Yao Hu,
  • JiaHong Yang,
  • YaQin Wang,
  • LiuMing Xiao

摘要

Multimodal action recognition, especially the fusion of image and skeleton data, has emerged as the prevailing approach in the field of action recognition. However, existing models often lack the spatial-temporal discriminative ability required for fine-grained recognition tasks. To address this issue, we introduce a flexible attention block called Variable-Channel Spatial-Temporal attention (VCSTA) to enhance the discriminative capacity of spatial-temporal connections. Based on VCSTA, we propose a novel multimodal variable-channel Spatial-Temporal semantic action recognition network (MMARN). MMARN utilizes a combination of spatial-temporal embedding loss and global loss to improve the model’s understanding of action changes and semantic information in videos, resulting in more precise prediction and analysis. The experimental results show that our multimodal variable-channel spatial-temporal semantic action recognition network achieves 98.3% and 89.9% accuracy in classifying actions on the large-scale human activity datasets NTU-RGB+D 60 and NTU-RGB+D 120 respectively.