Dynamic gesture recognition has indeed become a focal point in the field of human-computer interaction. Despite advancements, contemporary approaches fall short in considering the interdependence between frame images, and the complexity of real-world application backgrounds poses a significant obstacle to enhancing gesture recognition accuracy. In this paper, we propose a concatenated spatio-temporal attention with 3D convolutional network (CPTA3DNet) for gesture recognition. A stackable spatio-temporal attention module (STAM) is designed to effectively capture the dynamics of gestures over time and space. This module operates sequentially, initially emphasizing temporal attention followed by spatial attention. To evaluate the effectiveness of our method, we rigorously tested it on two large-scale public gesture recognition datasets: the Jester dataset and EgoGesture dataset. Focusing on the RGB modality, our experiments led to groundbreaking results, with our method attaining recognition accuracy of 94.79% on the Jester dataset and 94.36% on the EgoGesture dataset, which performs better than existing methods. Additionally, we performed comprehensive ablation studies to substantiate the impact of STAM in capturing both temporal and spatial dynamics.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention-Based Spatio-Temporal Modeling with 3D Convolutional Neural Networks for Dynamic Gesture Recognition

  • Yutong Hu

摘要

Dynamic gesture recognition has indeed become a focal point in the field of human-computer interaction. Despite advancements, contemporary approaches fall short in considering the interdependence between frame images, and the complexity of real-world application backgrounds poses a significant obstacle to enhancing gesture recognition accuracy. In this paper, we propose a concatenated spatio-temporal attention with 3D convolutional network (CPTA3DNet) for gesture recognition. A stackable spatio-temporal attention module (STAM) is designed to effectively capture the dynamics of gestures over time and space. This module operates sequentially, initially emphasizing temporal attention followed by spatial attention. To evaluate the effectiveness of our method, we rigorously tested it on two large-scale public gesture recognition datasets: the Jester dataset and EgoGesture dataset. Focusing on the RGB modality, our experiments led to groundbreaking results, with our method attaining recognition accuracy of 94.79% on the Jester dataset and 94.36% on the EgoGesture dataset, which performs better than existing methods. Additionally, we performed comprehensive ablation studies to substantiate the impact of STAM in capturing both temporal and spatial dynamics.