Activity Based Human Pose Estimation in Thermal Images
摘要
Human pose estimation using thermal images presents unique challenges, such as low resolution and varying heat signatures, which can make detecting body poses quite difficult. In this paper, we explore how deep learning can overcome these obstacles by fine-tuning pre-trained models like YOLOv8 and Vision Transformer. In this study, we compare and contrast various techniques for pose estimation on WICV Challenge on Thermal Human Pose Estimation, where we aim to improve pose detection accuracy in thermal imagery. By utilizing a publicly available dataset of annotated thermal images, we focus on achieving robust keypoint localization, even in complex real-world scenarios that involve diverse poses, occlusions, and different body shapes. Our experimental results demonstrate that the ViTPose++ model outperforms others in accuracy, proving it to be an excellent choice for applications where precision is critical. This research highlights the effectiveness of transformer-based architectures in thermal pose estimation, offering valuable insights for future advancements in analyzing thermal images.