Skeleton-based action recognition has received widespread attention in the field of video understanding. Recently, algorithms based on skeleton Gaussian pseudo-heatmaps have achieved excellent recognition performance due to their good robustness and action representation capabilities. However, these methods suffer from high computational costs. To solve these problems, we propose AM-SCNet, an efficient method to recognize actions by skeleton heatmaps. AM-SCNet applies sparse convolutions to ultimate feature extraction from skeleton pseudo heatmaps, significantly reducing redundant computations. We propose AM-RSC and AM-SSC by introducing Activation Mask (AM) into sparse convolutions for dynamically filtering spatiotemporal action semantics with significant sparse features. Experiments on NTU 60 and NTU 120 datasets have shown that AM-SCNet is highly efficient than. Compared to the advanced PoseC3D, our method achieves 93.8%, 96.7% accuracy on the X-Sub benchmark of NTU 60, while reducing GFLOPs by 85.9%. AM-SCNet offers an effective solution for achieving real-time skeleton-based action recognition tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AM-SCNet: Activation Mask Sparse Convolutional Network for Skeleton-Based Action Recognition

  • Yuzhou Gong,
  • Yunfei Xie,
  • Hanlin Li,
  • Yifan Han,
  • Huilin Ding,
  • Lan Luan,
  • Yuying Yong,
  • Shoudong Han

摘要

Skeleton-based action recognition has received widespread attention in the field of video understanding. Recently, algorithms based on skeleton Gaussian pseudo-heatmaps have achieved excellent recognition performance due to their good robustness and action representation capabilities. However, these methods suffer from high computational costs. To solve these problems, we propose AM-SCNet, an efficient method to recognize actions by skeleton heatmaps. AM-SCNet applies sparse convolutions to ultimate feature extraction from skeleton pseudo heatmaps, significantly reducing redundant computations. We propose AM-RSC and AM-SSC by introducing Activation Mask (AM) into sparse convolutions for dynamically filtering spatiotemporal action semantics with significant sparse features. Experiments on NTU 60 and NTU 120 datasets have shown that AM-SCNet is highly efficient than. Compared to the advanced PoseC3D, our method achieves 93.8%, 96.7% accuracy on the X-Sub benchmark of NTU 60, while reducing GFLOPs by 85.9%. AM-SCNet offers an effective solution for achieving real-time skeleton-based action recognition tasks.