Accurate reconstruction of 3D human pose from 2D keypoints is significantly enhanced by incorporating 3D bone length information. In this thesis, we design BTPose (Bone-to-Pose), a diffusion-based model that improves pose prediction by first estimating 3D bone lengths from either 2D pose inputs or direct measurements. These estimated bone lengths are then integrated with the 2D pose to produce multiple plausible 3D pose candidates. To additionally enhance the accuracy of estimating 3D human poses, we introduce N-JPMA, an enhanced version of the Joint-wise reProjection-based Multi-hypothesis Aggregation (JPMA) method. Unlike the JPMA method, which selects a single 3D pose whose projected result best matches the given 2D observations, N-JPMA averages the top-N hypotheses with the smallest re-projection errors to obtain the final prediction, effectively mitigating challenges such as unreliable 2D poses and the absence of depth information. Extensive experiments on commonly adopted datasets, including Human3.6M and MPI-INF-3DHP, demonstrate that BTPose outperforms top-performing methods, including both disentangled and non-disentangled models, as well as probabilistic approaches, considering both accuracy and robustness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BTPose: 3D Pose Estimation from Bone to Pose with Efficient Multi-hypothesis Aggregation

  • Jingtian Li,
  • Yi Wu,
  • Shangfei Wang,
  • Guoming Li,
  • Meng Mao,
  • Linxiang Tan

摘要

Accurate reconstruction of 3D human pose from 2D keypoints is significantly enhanced by incorporating 3D bone length information. In this thesis, we design BTPose (Bone-to-Pose), a diffusion-based model that improves pose prediction by first estimating 3D bone lengths from either 2D pose inputs or direct measurements. These estimated bone lengths are then integrated with the 2D pose to produce multiple plausible 3D pose candidates. To additionally enhance the accuracy of estimating 3D human poses, we introduce N-JPMA, an enhanced version of the Joint-wise reProjection-based Multi-hypothesis Aggregation (JPMA) method. Unlike the JPMA method, which selects a single 3D pose whose projected result best matches the given 2D observations, N-JPMA averages the top-N hypotheses with the smallest re-projection errors to obtain the final prediction, effectively mitigating challenges such as unreliable 2D poses and the absence of depth information. Extensive experiments on commonly adopted datasets, including Human3.6M and MPI-INF-3DHP, demonstrate that BTPose outperforms top-performing methods, including both disentangled and non-disentangled models, as well as probabilistic approaches, considering both accuracy and robustness.