This paper presents a novel approach for human action recognition in videos by focusing on the extraction of spatio-temporal features. Recognizing actions in videos is of paramount importance for various applications, including video surveillance, sports analysis, and human-computer interaction. To address this challenge, we proposed an hybrid network that integrates a time-distributed wrapper and attention-based mechanisms. The incorporation of a time-distributed wrapper enables the model to effectively extract the temporal dynamics of actions by extending the capabilities of Convolutional Neural Networks (CNNs) to process sequences of frames. Moreover, the integration of attention-based mechanisms allows the network to selectively weigh and emphasize relevant spatio-temporal features, further enhancing its discriminative power. The proposed network outperforms the state-of-the-art methods in human action recognition, achieving an impressive average accuracy of 99.3% and 99.49% on the UCF101 and UCF11 datasets respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Video Surveillance: Synergizing Time-Distributor Wrapper and Attention Mechanisms for Superior Human Action Recognition

  • Khawla Ben Salah,
  • Mohamed Othmani,
  • Jihen Fourati,
  • Monji Kherallah

摘要

This paper presents a novel approach for human action recognition in videos by focusing on the extraction of spatio-temporal features. Recognizing actions in videos is of paramount importance for various applications, including video surveillance, sports analysis, and human-computer interaction. To address this challenge, we proposed an hybrid network that integrates a time-distributed wrapper and attention-based mechanisms. The incorporation of a time-distributed wrapper enables the model to effectively extract the temporal dynamics of actions by extending the capabilities of Convolutional Neural Networks (CNNs) to process sequences of frames. Moreover, the integration of attention-based mechanisms allows the network to selectively weigh and emphasize relevant spatio-temporal features, further enhancing its discriminative power. The proposed network outperforms the state-of-the-art methods in human action recognition, achieving an impressive average accuracy of 99.3% and 99.49% on the UCF101 and UCF11 datasets respectively.