Deep learning based on hand pose estimation methods: a systematic literature review
摘要
Estimating hand pose is a challenge that has significantly benefited from using deep learning-based algorithms. This study area holds critical significance across various computer vision and robotics domains, including applications in sign language interpretation, Computer-Aided Design (CAD), and 3D humanoid reconstruction systems. These technologies flourish in augmented reality systems, facilitating immersive interactions within virtual reality contexts. The complexity of human hand anatomy and its intricate range of motions intensifies the difficulty of accurate pose estimation, posing significant academic and technical hurdles. Recent advancements have seen the emergence of rapid and comprehensive methods for hand pose estimation, driven by advancements in-depth camera technology and Deep Neural Networks (DNNs). The previous surveys studied most of these hand pose issues, including hand parsing, data labeling technologies, hand motion, fingertip detection, hand localization, and self-occlusion. This paper addresses the aforementioned challenges in hand pose estimation. Further, we propose a novel taxonomy based on deep-learning-based approaches that aim to gather the previous research advances systems that tackled these actual challenges. We provide an overview of existing research, discussing their strengths and limitations. Additionally, we identify various benchmark datasets, their characteristics and prevalent evaluation metrics used to assess these approaches. Finally, we explore potential research directions focusing on speed, accuracy, and type of deep learning architecture in this rapidly evolving field.