3D hand pose estimation based on lightweight CNN and separable self-attention vision transformer
摘要
At present, most 3D hand pose estimation methods using RGB images suffer from high computational complexity and slow inference. To address these issues, we propose a method titled “3D Hand Pose Estimation Based on Lightweight CNN and Separable Self-Attention Vision Transformer”, which leverages lightweight convolutional neural networks and separable self-attention Vision Transformers. The backbone network integrates Sandglass block to efficiently extract local features, with Vision Transformers incorporating separable self-attention to enhance global feature extraction. Compared to mainstream methods, the 3D hand pose estimation method achieves higher accuracy and significantly faster processing speed on the RHD and STB datasets. This method achieves an optimal balance between accuracy and speed, crucial for real-time applications. The proposed method substantially enhances 3D hand pose estimation by effectively addressing the challenges of computational complexity and slow inference.