Unsupervised Feature Learning for Video Understanding
摘要
Video understanding with deep learning has been advanced by leaps and bounds by virtue of both improvements in deep neural networks and the emergence of large-scale training datasets. Vast amounts of annotated data have led to the growth in the performance of supervised learning; nevertheless, manual collection and annotation are demanding of time and labor. Subsequently, research interests have been aroused in unsupervised feature learning that indicates learning representations from data that exempt from human supervision. In the meantime, deep learning on videos faced more challenges than images due to the intricacy of modeling temporal information. To address such an issue, increasing effort has been made in tackling video representation with unlabeled training samples, referred to as unsupervised video learning (UVL). Particularly, the intrinsic temporal context in videos serves as the hotbeds of pretext tasks, which refer to network optimization tasks based on surrogate signals without human supervision, facilitating better performance on video-related downstream tasks. In this chapter, we undertake a comprehensive review of UVL, which begins with a preliminary introduction of unsupervised approaches and proceeds with the elaboration of self-supervised video learning (SSVL) driven by different pretext tasks. In particular, SSVL refers to a typical category of unsupervised techniques, where the supervisory signals are mined from videos.