TSGFormer: temporal-aware network and spatial encoding GCN for three-dimensional human pose estimation
摘要
Transformer-based approaches have significantly driven recent progress in three-dimensional human pose estimation. However, existing transformer-based approaches are still deficient in capturing localized features, and they lack task-specific a priori information by obtaining queries, keys, and values through simple linear mappings. Existing methods lack effective human constraints for model training. We introduce the Spatial Encoding Graph Convolutional Network Transformer (SEGCNFormer), designed to enhance model capacity in capturing local features. In addition, we propose a Temporal-Aware Network, which generates queries, keys, and values possessing a priori knowledge of human motion, enabling the model to better understand the structural information of human poses. Finally, we leverage the knowledge of human anatomy and motion to design the Human Structural Science Loss, which performs a rationality assessment of human actions and imposes physical constraints on the generated poses. Our method outperforms existing methods on the Human3.6M dataset in both 27 and 81 sampling frames, and our predicted poses are closer to the actual poses with less error. For the existing three issues, we proposed effective methods and conducted targeted experiments, which confirmed the effectiveness of our strategies.