ST-HViT: spatial-temporal hierarchical vision transformer for action recognition
摘要
Human action recognition (HAR) is an important task in the field of computer vision, and its primary goal is to analyze and understand human activities in videos. In addition to containing the spatial information of static images, videos also contain unique temporal information, which makes the information contained in the videos even richer. However, the training cost required to fully learn the spatial-temporal information of the videos is quite expensive for a model. In light of this, we propose a novel two-stream network structure to effectively capture the spatial-temporal information in video data. We perform masked autoencoders (MAE) pre-training, aiming to reduce the training burden of the model through this asymmetric encoder-decoder pre-training method. In addition, we propose a new multi-scale decoder component that combines transposed convolutional upsampling and convolutional downsampling. It fully utilizes the multi-scale features of the encoder to achieve excellent performance. On two challenging video datasets, Kinetics 400 (K400) and Something-Something-v2 (SSv2), we achieve state-of-the-art performance with 85.9