Graph Convolutional Networks (GCNs) have shown remarkable success in 3D human pose estimation. Nonetheless, effectively modeling the motion correlation among human joints remains a challenging task. This paper presents a novel Dynamic Learning Specific Motion Graph Convolutional Network (DLSM-GCN) block, aimed at adaptively capturing the kinematic relationships between human joints from diverse inputs. Specifically, it introduces two components: Dynamic Spatio-Temporal Graph Convolution (DSTG) and Dynamic Spatial Second-Order Connectivity Graph Convolution (DSOG). DSTG is devised to model latent motion correlations between joints, while DSOG focuses on capturing second-order connectivity relationships among joints exhibiting vigorous motion. Furthermore, we propose the DSLMFormer, which integrates the DLSM-GCN block with the Temporal Transformer Block (TTB). DLSM-GCN and TTB respectively handle spatial and temporal modeling of human pose in videos, effectively mitigating depth ambiguity and motion uncertainty. Extensive experiments conducted on various datasets demonstrate that DLSMFormer dynamically captures motion correlations among human joints from diverse inputs, thereby enhancing the accuracy of 3D human pose estimation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning the Dynamic Spatio-Temporal Relationship Between Joints for 3D Human Pose Estimation

  • Feiyi Xu,
  • Ying Sun,
  • Jin Qi,
  • Yanfei Sun

摘要

Graph Convolutional Networks (GCNs) have shown remarkable success in 3D human pose estimation. Nonetheless, effectively modeling the motion correlation among human joints remains a challenging task. This paper presents a novel Dynamic Learning Specific Motion Graph Convolutional Network (DLSM-GCN) block, aimed at adaptively capturing the kinematic relationships between human joints from diverse inputs. Specifically, it introduces two components: Dynamic Spatio-Temporal Graph Convolution (DSTG) and Dynamic Spatial Second-Order Connectivity Graph Convolution (DSOG). DSTG is devised to model latent motion correlations between joints, while DSOG focuses on capturing second-order connectivity relationships among joints exhibiting vigorous motion. Furthermore, we propose the DSLMFormer, which integrates the DLSM-GCN block with the Temporal Transformer Block (TTB). DLSM-GCN and TTB respectively handle spatial and temporal modeling of human pose in videos, effectively mitigating depth ambiguity and motion uncertainty. Extensive experiments conducted on various datasets demonstrate that DLSMFormer dynamically captures motion correlations among human joints from diverse inputs, thereby enhancing the accuracy of 3D human pose estimation.