An in-depth exploration of structural pose estimation strategies and datasets
摘要
Human Pose Estimation (HPE) refers to detecting the major joints (skeleton) of a human and their motion in order to recognize any action. This task is important because of its serious implications, such as surveillance, autonomous vehicles, sports analytics, human-computer applications, animation, and medical tracking. This article discusses various methods used to survey HPE, ranging from the fundamental principles of computer vision to advanced deep learning (DL) models. We investigate various structural pose estimation strategies for 2D and 3D domains and explore volumetric, planar, and kinematic models. Not only does this paper investigate methodology, but we also evaluate popular benchmark datasets used to test these methods, looking into their pros, cons, and overall usefulness for various HPE problems. Depth ambiguity, viewpoint variation, occlusion, and the absence of any annotated data are described as some of the many critical challenges to overcome. We further investigate what can be done to assist in the use of robotics and surveillance in real-time, such as self-supervised learning and adapting to specific task domains, aiming to drive future efforts forward.