Attention-Based Spatio-Temporal Modeling with 3D Convolutional Neural Networks for Dynamic Gesture Recognition
摘要
Dynamic gesture recognition has indeed become a focal point in the field of human-computer interaction. Despite advancements, contemporary approaches fall short in considering the interdependence between frame images, and the complexity of real-world application backgrounds poses a significant obstacle to enhancing gesture recognition accuracy. In this paper, we propose a concatenated spatio-temporal attention with 3D convolutional network (CPTA3DNet) for gesture recognition. A stackable spatio-temporal attention module (STAM) is designed to effectively capture the dynamics of gestures over time and space. This module operates sequentially, initially emphasizing temporal attention followed by spatial attention. To evaluate the effectiveness of our method, we rigorously tested it on two large-scale public gesture recognition datasets: the Jester dataset and EgoGesture dataset. Focusing on the RGB modality, our experiments led to groundbreaking results, with our method attaining recognition accuracy of 94.79% on the Jester dataset and 94.36% on the EgoGesture dataset, which performs better than existing methods. Additionally, we performed comprehensive ablation studies to substantiate the impact of STAM in capturing both temporal and spatial dynamics.