Exploring the Impact of Convolutions on LSTM Networks for Video Classification
摘要
Video classification plays a foundational role within the field of computer vision, that involves categorizing and labeling videos based on their content. Its significance is evident in a wide array of applications, encompassing video surveillance, content recommendation, action recognition, video indexing, and more. The goal of video classification is to automatically analyze and understand the visual information present in videos, enabling efficient organization, retrieval, and interpretation of large video collections. The fusion of convolutional neural networks (CNNs) and long short term memory (LSTM) networks has revolutionized the field of video classification by effectively capturing both spatial and temporal dependencies within video sequences. This fusion combines the strengths of CNNs in extracting spatial features and LSTMs in modeling sequential and temporal information. Two widely adopted architectures that incorporate this fusion are ConvLSTM and LRCN (Long-term Recurrent Convolutional Networks). This paper aims to explore the impact of convolutions on LSTM networks in the context of video classification and compare the performance of ConvLSTM and LRCN.