Action recognition, a crucial branch of computer vision, exhibits vast application potential in human computer interaction, video surveillance, and more. The core challenge lies in accurately modeling the action information within videos, primarily involving three key elements: spatio-temporal information, channel information, and motion information. Traditional convolutional neural networks (CNNs) suffer from limitations in temporal modeling, directly impacting recognition accuracy. Recently emerged video Transformer networks, while overcoming this limitation to some extent with their inherent global temporal modeling capabilities, come with high computational and training costs. To address these issues, this paper proposes a novel action recognition model based on the deep fusion of spatio-temporal, channel, and motion features. Specifically, a Motion Feature Recognition (MFR) module is designed. It employs a dual-differencing operation on feature maps, effectively extracting motion features while significantly reducing redundancy in the difference maps. Furthermore, we introduce the Channel Feature Recognition (CFR) module and a Spatial-Temporal (S-T) module, which work collaboratively to enable the model to more accurately capture channel features containing key action information and the most important spatio-temporal features. To validate the model’s effectiveness, we conducted comprehensive comparative experiments on three authoritative datasets: UCF101 action Dataset, Something-Something-V1, and Something-Something-V2. Experimental results show that compared to CNN-based action recognition models, our proposed model achieves a significant improvement in recognition accuracy. Simultaneously, while maintaining performance comparable to Transformer frameworks, our model significantly reduces FLOPs (floating-point operations), demonstrating higher computational efficiency. Further ablation experiments strongly demonstrate the complementarity of spatio-temporal information, channel information, and motion information, as well as the significant effect of their fusion, thus fully validating the effectiveness and advancement of our proposed method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MCSTN: A Novel Multi-feature Fusion-Based Action Recognition Model

  • Wei Wei,
  • Ang Li,
  • Xiaodong Duan

摘要

Action recognition, a crucial branch of computer vision, exhibits vast application potential in human computer interaction, video surveillance, and more. The core challenge lies in accurately modeling the action information within videos, primarily involving three key elements: spatio-temporal information, channel information, and motion information. Traditional convolutional neural networks (CNNs) suffer from limitations in temporal modeling, directly impacting recognition accuracy. Recently emerged video Transformer networks, while overcoming this limitation to some extent with their inherent global temporal modeling capabilities, come with high computational and training costs. To address these issues, this paper proposes a novel action recognition model based on the deep fusion of spatio-temporal, channel, and motion features. Specifically, a Motion Feature Recognition (MFR) module is designed. It employs a dual-differencing operation on feature maps, effectively extracting motion features while significantly reducing redundancy in the difference maps. Furthermore, we introduce the Channel Feature Recognition (CFR) module and a Spatial-Temporal (S-T) module, which work collaboratively to enable the model to more accurately capture channel features containing key action information and the most important spatio-temporal features. To validate the model’s effectiveness, we conducted comprehensive comparative experiments on three authoritative datasets: UCF101 action Dataset, Something-Something-V1, and Something-Something-V2. Experimental results show that compared to CNN-based action recognition models, our proposed model achieves a significant improvement in recognition accuracy. Simultaneously, while maintaining performance comparable to Transformer frameworks, our model significantly reduces FLOPs (floating-point operations), demonstrating higher computational efficiency. Further ablation experiments strongly demonstrate the complementarity of spatio-temporal information, channel information, and motion information, as well as the significant effect of their fusion, thus fully validating the effectiveness and advancement of our proposed method.