Capturing discriminative spatiotemporal feature information is crucial for enhancing the performance of video action recognition. Currently, most methods generally embed spatiotemporal feature modeling modules into the backbone network to achieve human action recognition under a single view. However, action video has a 3D structure, and multi-view information helps to acquire more comprehensive body-related clues to improve the robustness and accuracy of action recognition. Depth data is insensitive to lighting and color changes, particularly in providing reliable 3D geometric information of human actions. In this paper, we focus on the task of depth video-based action recognition and propose a Dual-view Spatio-Temporal Interactive Network (DSTIN), which enhances action recognition performance by effectively integrating spatiotemporal interactive features from front and side views. Specifically, the core of the DSTIN is the Dual-view Spatio-Temporal Interactive Module (DSTIM), which operates on convolutional feature maps from two views. It first exchanges spatiotemporal information along the channel dimension to achieve inter-view channel feature interaction, and then applies temporal shift along the time dimension within each view to model motion information. The proposed DSTIM can be flexibly embedded into any 2D deep network architectures to enhance their ability. We choose 2D Resnet50 as the backbone to generate DSTIN, which takes depth dynamic image sequences from two different views as the input and achieves an end-to-end action recognition training framework through inter-view spatiotemporal information interaction and fusion. Extensive experiments on two large-scale RGBD datasets demonstrate that the proposed method significantly improves the performance of video human action recognition.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual-View Spatio-Temporal Interactive Network for Video Human Action Recognition

  • Hanbo Wu,
  • Xin Ma

摘要

Capturing discriminative spatiotemporal feature information is crucial for enhancing the performance of video action recognition. Currently, most methods generally embed spatiotemporal feature modeling modules into the backbone network to achieve human action recognition under a single view. However, action video has a 3D structure, and multi-view information helps to acquire more comprehensive body-related clues to improve the robustness and accuracy of action recognition. Depth data is insensitive to lighting and color changes, particularly in providing reliable 3D geometric information of human actions. In this paper, we focus on the task of depth video-based action recognition and propose a Dual-view Spatio-Temporal Interactive Network (DSTIN), which enhances action recognition performance by effectively integrating spatiotemporal interactive features from front and side views. Specifically, the core of the DSTIN is the Dual-view Spatio-Temporal Interactive Module (DSTIM), which operates on convolutional feature maps from two views. It first exchanges spatiotemporal information along the channel dimension to achieve inter-view channel feature interaction, and then applies temporal shift along the time dimension within each view to model motion information. The proposed DSTIM can be flexibly embedded into any 2D deep network architectures to enhance their ability. We choose 2D Resnet50 as the backbone to generate DSTIN, which takes depth dynamic image sequences from two different views as the input and achieves an end-to-end action recognition training framework through inter-view spatiotemporal information interaction and fusion. Extensive experiments on two large-scale RGBD datasets demonstrate that the proposed method significantly improves the performance of video human action recognition.