Compared to multi-view 3D human pose estimation, monocular 3D human pose estimation has more important applications in human-computer interaction and human action recognition. Simultaneously achieving real-time speed, varying human numbers, and high accuracy from a single RGB image are challenging problems. To this end, this chapter proposes a multitask and multi-level neural network structure with physical constraints. The unique network structure estimates 3D human poses from a single RGB image in an end-to-end way and achieves both high accuracy and high speed. Experimental results show that the proposed system achieves 21 FPS on RTX 2080 GPU with only 33 mm accuracy loss compared with conventional works. The mechanism of the network is also analyzed through network visualization. This work shows the possibility of estimating 3D human pose from a single RGB monocular camera with real-time speed.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Real-Time 3D Human Pose Estimation from a Single RGB Image

  • Songlin Du,
  • Takeshi Ikenaga

摘要

Compared to multi-view 3D human pose estimation, monocular 3D human pose estimation has more important applications in human-computer interaction and human action recognition. Simultaneously achieving real-time speed, varying human numbers, and high accuracy from a single RGB image are challenging problems. To this end, this chapter proposes a multitask and multi-level neural network structure with physical constraints. The unique network structure estimates 3D human poses from a single RGB image in an end-to-end way and achieves both high accuracy and high speed. Experimental results show that the proposed system achieves 21 FPS on RTX 2080 GPU with only 33 mm accuracy loss compared with conventional works. The mechanism of the network is also analyzed through network visualization. This work shows the possibility of estimating 3D human pose from a single RGB monocular camera with real-time speed.