Introduction
摘要
With universal access to cameras in all kinds of devices (e.g., mobile phones, surveillance cameras, in-vehicle cameras), video data has gone through an exponential increase nowadays. The ability to understand, track, and segment objects in videos has become a fundamental and essential problem in various video-related applications. For example, in the field of video and movie editing, it is a very common task to separate the pixels of a foreground apart from the pixels of the background in the original video, and then the foreground pixels are put onto some new background to create fancy visual effects.