Infant Video Interaction Recognition Using Monocular Depth Estimation
摘要
We present a system for automatic analysis of videos of infant actions and infant-caregiver interactions during in-home play sessions. Variables examined include the posture of the infant (sitting, standing, prone, supine, etc.) and whether they are supported in that position by an inanimate object or assisted by a caregiver and at what body location. Leveraging recent advances in neural monocular depth estimation, human body keypoints are lifted from 2-D to 3-D to compute metric distance and angle features, and 3-D scene properties such as the floor plane are estimated to put detections in a global spatial context. We demonstrate strong performance on related pose estimation and posture benchmarks as well as vs. state-of-the-art methods on a challenging new naturalistic video dataset featuring complex interactions in cluttered scenes. We believe that this approach shows promise as a tool for scaling up infant motor developmental studies and is extensible to other developmental domains and age groups.