Video-Based Self-supervised Human Depth Estimation
摘要
In this paper, we propose a video-besed method for self-supervised human depth estimation, aiming at the problem of joint point distortion in human depth and insufficient utilization of 3D information in video-based depth estimation. We use the relative ordinal relations between human joint point pairs to deal with the problem of joint point distortion. Meanwhile, a temporal correlation module is proposed to focus on the temporal correlation between past and present frames, taking into account the influence of temporal characteristics in the video sequence. A hierarchical structure is adopted to fuse adjacent features, thus fully mine the 3D information based on the video. The experimental results show that this model significantly improves the human depth estimation performance, especially at the joints.