2D human pose estimation from thermal infrared images has remained an understudied yet crucial research problem, largely due to the scarcity of annotated data, which hinders the development of accurate models. To overcome this limitation, we propose a novel unsupervised training paradigm that leverages domain transformation to bridge the gap between thermal and RGB modalities. Specifically, we apply a domain transform to thermal images, converting them into RGB-like representations that can be seamlessly integrated with existing state-of-the-art 2D pose estimation methods. The resulting pose points, generated from these transformed images, serve as pseudo Ground Truth (GT) labels to train a thermal pose estimation encoder-decoder network. Furthermore, we extend this network to a multi-modal framework, where the same encoder is shared across both pose estimation and activity recognition tasks, with separate decoders for each task. Our approach exploits the complementary information between these two tasks, demonstrating that the optimization of the encoder network for activity recognition yields robust features that, in turn, improve the performance of the 2D pose estimation decoder. Experimental results on the LWIRPOSE dataset [23] demonstrate that our unsupervised methodology achieves comparable performance to supervised approaches, highlighting its potential for diverse applications in thermal imaging.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unsupervised Thermal Image Pose Estimation and Activity Recognition via Thermal to RGB Domain Transform and Multi-Task Learning

  • Avinash Upadhyay,
  • Bhipanshu Dhupar,
  • Manoj Sharma,
  • Shreya Sharma,
  • Nikhil Kumar,
  • Santanu Chaudhury

摘要

2D human pose estimation from thermal infrared images has remained an understudied yet crucial research problem, largely due to the scarcity of annotated data, which hinders the development of accurate models. To overcome this limitation, we propose a novel unsupervised training paradigm that leverages domain transformation to bridge the gap between thermal and RGB modalities. Specifically, we apply a domain transform to thermal images, converting them into RGB-like representations that can be seamlessly integrated with existing state-of-the-art 2D pose estimation methods. The resulting pose points, generated from these transformed images, serve as pseudo Ground Truth (GT) labels to train a thermal pose estimation encoder-decoder network. Furthermore, we extend this network to a multi-modal framework, where the same encoder is shared across both pose estimation and activity recognition tasks, with separate decoders for each task. Our approach exploits the complementary information between these two tasks, demonstrating that the optimization of the encoder network for activity recognition yields robust features that, in turn, improve the performance of the 2D pose estimation decoder. Experimental results on the LWIRPOSE dataset [23] demonstrate that our unsupervised methodology achieves comparable performance to supervised approaches, highlighting its potential for diverse applications in thermal imaging.