Dynamic Gesture Recognition Based on 3D Central Difference Separable Residual LSTM Coordinate Attention Networks
摘要
The recognition of dynamic gestures has garnered significant attention in the field of human-computer interaction. However, several factors unrelated to the gestures, such as background, and spatial scale, pose significant challenges in improving the accuracy of recognition, despite the flexibility of the gestures themselves. In this paper, we propose an end-to-end recognition network, named 3D central difference separable residual long and short-term memory (LSTM) coordinate attention (3D CRLCA), which addresses these issues. Our network utilizes 3D central difference separable convolution (3D CDSC) to extract fine-grained spatiotemporal feature information, facilitating the recognition and classification of dynamic gestures. We have also incorporated a residual module in the network to enhance the discriminative ability between gesture categories. To further extract semantic and action information of the gestures, we have combined the LSTM-CA attention mechanism, which enables the network to focus on the gesture area and the temporal and spatial characteristics of gestures, to facilitate gesture recognition. Our experiments on the ChaLearn Large-scale Gesture Recognition Dataset (IsoGD) and IPN dataset demonstrate that our approach outperforms other methods.