Apart from gesture recognition and augmented reality applications, pose estimation has been utilized in the medical field to aid clinicians in conducting patient care. Existing 3D hand pose estimation methods typically adopt deep learning networks that exhibit high computational complexity, making the deployment of these methods into mobile devices difficult. To address this problem, we propose a novel knowledge distillation framework that incorporates hand kinematics through segment distance and segment direction. This maximizes the knowledge gained by the student network while maintaining its size. Experimental results on the FreiHAND dataset using different variations of the framework demonstrate how student network performance can be improved by incorporating segment distance and segment direction into the knowledge distillation process. The framework proposed in this work produced a lightweight network with 68% fewer trainable parameters than the state-of-the-art network.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Lightweight 3D Hand Pose Estimation Using Knowledge Distillation with Hand Kinematics

  • Jose Lorenzo C. Capistrano,
  • Nathaniel S. Orillaza,
  • Prospero C. Naval

摘要

Apart from gesture recognition and augmented reality applications, pose estimation has been utilized in the medical field to aid clinicians in conducting patient care. Existing 3D hand pose estimation methods typically adopt deep learning networks that exhibit high computational complexity, making the deployment of these methods into mobile devices difficult. To address this problem, we propose a novel knowledge distillation framework that incorporates hand kinematics through segment distance and segment direction. This maximizes the knowledge gained by the student network while maintaining its size. Experimental results on the FreiHAND dataset using different variations of the framework demonstrate how student network performance can be improved by incorporating segment distance and segment direction into the knowledge distillation process. The framework proposed in this work produced a lightweight network with 68% fewer trainable parameters than the state-of-the-art network.