DT2S-Pose: A Deeper Temporal-Spatial Skeleton Refine Model for Pedestrian Pose Estimation
摘要
Task-specific multi-frame human pose estimation is challenging. While state-of-the-art multi-frame human keypoints estimation techniques have achieved significant results on general datasets, estimating keypoints for specific scenarios still needs to be improved (e.g., walking scenarios). The main drawbacks include the inability to distinguish similar parts, posture occlusion, motion blur, etc. This is due to the inability of existing network structures to combine spatial and temporal contexts thoroughly. In this paper, we propose a new multi-frame human pose estimation model for gait tasks, which fully considers contextual information of the temporal and spatial layers to detect keypoints in fixed scenes. Our framework designs two modular components: the Temporal-Spatial Refine Transformer Mixture (TSRTM) module that jointly learns spatial-temporal context at a deeper level and the Skeleton Correct Network (SCN) that considers the human physiological structure and leverages the relationships between multiple points. The keypoint annotations about the partial CASIA-B dataset, which is a large-scale dataset designed for gait recognition tasks, are proposed in this paper. We improve the research of the downstream task by designing human pose estimation models for walking groups. Our method performs well on the CASIA-B dataset and is robust under occlusion conditions. We also demonstrate the generality of our method on the PoseTrack2017 dataset.