Hybrid attention-inflated 3D architecture for human action recognition
摘要
Human Action Recognition (HAR) in videos is a complex challenge primarily due to the difficulty of simultaneously capturing and integrating spatial and temporal information from video sequences. Traditional methods often struggle with this duality, leading to suboptimal performance in recognizing intricate human actions. To address these limitations, we propose a novel approach that effectively integrates both spatial and temporal features using a hybrid architecture combining the Convolutional Block Attention Module (CBAM) and Inflated 3D Convolutional Networks (I3D), referred to as CBAM-I3D. Our approach consists of two key components: (1) The CBAM attention block, which enhances feature extraction by focusing on both channel-wise and spatial-temporal attention, thereby improving the representation of crucial spatio-temporal features. (2) The Inflated 3D ConvNet (I3D), which extends conventional 2D convolutions into 3D to capture dynamic motion patterns across frames. The synergy between these components allows for a more robust and accurate recognition of human activities by leveraging enriched feature representations. The obtained results show that our proposed model achieves good performance on two of the most challenging datasets including UCF-101 (92.42%) and HMDB-51 (91.0%).