Research on 3D human pose estimation via attention-guided adjacent frame-aware network
摘要
Existing monocular video-based 3D pose estimation techniques often fail to fully exploit dynamic information across consecutive frames, resulting in incoherent pose estimations and significant degradation in joint localization accuracy under rapid motions or occlusions. To address these challenges, this paper proposes a novel Attention-guided Adjacent Frame-aware Network (AAFN) for 3D human pose estimation. The AAFN incorporates a Guided Temporal Attention Module (GTAM) to capture temporal dependencies and an External Guided Temporal Focus Module (EGTF) to enhance keyframe features. Based on the GTAM framework, we design a Neighborhood-Aware Temporal Encoder (NATE) to fuse local details from adjacent frames, coupled with bidirectional GRUs for feature smoothing, thereby constructing refined spatiotemporal representations. Experimental results demonstrate that the AAFN achieves MPJPE and PA-MPJPE scores of 63.8 mm and 42.2 mm, respectively, on the Human3.6M dataset, outperforming the baseline model TCMR by 13.3% and 18.8%.