错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Resstanet: deep residual spatio-temporal attention network for violent action recognition

  • Ajeet Pandey,
  • Piyush Kumar

摘要

Violent Action Recognition (VAR) is a critical domain of research within computer vision and artificial intelligence, aiming to automatically detect violent behaviors in videos. Most of the current VAR methods are not able to extract spatial and temporal features simultaneously, which is crucial for human action recognition. This paper introduces a unique Residual Spatio-Temporal Attention Network (ResSTANet) model for robust violent action recognition. The ResSTANet uses a residual 3-Dimensional Convolutional Neural Network (3D-CNN) for effectively capturing spatiotemporal dynamics simultaneously. A residual connection is incorporated to improve information flow and control critical spatial-temporal features. The output of the residual 3D-CNN is subjected to a Multi-Head Attention (Mu-HA) process, increasing the focus on crucial features. Subsequently, multiple dense and dropout layers are applied to refine feature selection and reduce noise in the representation. Finally, a softmax layer is applied to perform action recognition, achieving state-of-the-art performance in VAR tasks.