Exploring CNN-Based Algorithms for Human Action Recognition in Videos
摘要
This study presents a comparative analysis of three convolutional neural network (CNN)-based methodologies, namely the Two-Stream CNN, CNN + LSTM, and 3D CNN, for human action recognition in video sequences. The main goal of this research is to analyze and understand human behaviors in video content. And subsequently, generate associated tags, all while surmounting the intricate spatial and temporal intricacies inherent in this task. The experimental evaluation employs the HMDB-51 dataset, and the findings reveal that all three proposed algorithms effectively discern human actions within the video domain, albeit with distinct performance variations. Furthermore, the paper offers in-depth elucidations and comprehensive analyses of each of these methods, thereby imparting valuable insights and directions for prospective research endeavors in the realm of human action recognition.