GCTT: Graph Convolution and Time-Frequency Integration Network for 3D Human Pose Estimation
摘要
Recently, methods based on graph convolution have shown significant advancements in the domain of 2D to 3D human pose estimation. However, the majority of these approaches primarily concentrate on spatial feature extraction from 2D poses, neglecting the comprehensive utilization of temporal information embedded within the pose sequences. Consequently, this paper proposes a novel 3D human pose estimation network named the Graph Convolution and Temporal Transformer Network (GCTT). The GCTT model extracts both spatial features and temporal features from human poses. During evaluation on the Human3.6M dataset, the Mean Per Joint Position Error (MPJPE) attained significant values of 33.9 mm.