Deep Learning for Action Recognition
摘要
The fast development of deep learning has significantly advanced video understanding, facilitating a shift from traditional manual feature-based approaches to automated analysis strategies based on deep neural networks. In this chapter, we review the evolution of deep learning methods for video content classification, which includes categorizing human activities and complex events in videos. We provide a detailed discussion on Convolutional Neural Network (CNN)-based and Transformer-based deep learning algorithms for video classification together with long-range temporal feature aggregation methods. Furthermore, we present an overview of commonly used video action recognition datasets, demonstrating how increasingly challenging datasets motivate the development of video understanding techniques.