<p>At present, most 3D hand pose estimation methods using RGB images suffer from high computational complexity and slow inference. To address these issues, we propose a method titled <i>“3D Hand Pose Estimation Based on Lightweight CNN and Separable Self-Attention Vision Transformer”</i>, which leverages lightweight convolutional neural networks and separable self-attention Vision Transformers. The backbone network integrates Sandglass block to efficiently extract local features, with Vision Transformers incorporating separable self-attention to enhance global feature extraction. Compared to mainstream methods, the 3D hand pose estimation method achieves higher accuracy and significantly faster processing speed on the RHD and STB datasets. This method achieves an optimal balance between accuracy and speed, crucial for real-time applications. The proposed method substantially enhances 3D hand pose estimation by effectively addressing the challenges of computational complexity and slow inference.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

3D hand pose estimation based on lightweight CNN and separable self-attention vision transformer

  • Tianpei Jin,
  • Qingshan She,
  • Qixiang Wang,
  • Qiang Chen,
  • Xiaofei Zhou

摘要

At present, most 3D hand pose estimation methods using RGB images suffer from high computational complexity and slow inference. To address these issues, we propose a method titled “3D Hand Pose Estimation Based on Lightweight CNN and Separable Self-Attention Vision Transformer”, which leverages lightweight convolutional neural networks and separable self-attention Vision Transformers. The backbone network integrates Sandglass block to efficiently extract local features, with Vision Transformers incorporating separable self-attention to enhance global feature extraction. Compared to mainstream methods, the 3D hand pose estimation method achieves higher accuracy and significantly faster processing speed on the RHD and STB datasets. This method achieves an optimal balance between accuracy and speed, crucial for real-time applications. The proposed method substantially enhances 3D hand pose estimation by effectively addressing the challenges of computational complexity and slow inference.