In this paper, we present a novel yet intuitive unsupervised feature learning approach, referred to as Minimizing Interframe Differences (MID). The idea is the following: as long as the unsupervised features successfully encode the essential information about the visual structures of the frames, the differences between the before and after frames can be minimally encoded. MID is implemented with a difference encoding module and a frame reconstruction module. The former tries to minimize the difference between frames, while the latter first maps the difference encoding to optical flow and then realizes frame reconstruction based on the learned optical flow. All the modules are jointly learned in an end-to-end way through the reconstruction loss. We verify the feasibility of MID in the image datasets by replacing the natural transformation of videos to artificially parameterized transformation for images. Experimental results show that MID is able to predict optical flow accurately based on the minimum difference coding and learn a powerful visual representation that greatly enhances other visual tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning Features by Minimizing the Interframe Differences

  • Dong Zhao,
  • Dan Zhang

摘要

In this paper, we present a novel yet intuitive unsupervised feature learning approach, referred to as Minimizing Interframe Differences (MID). The idea is the following: as long as the unsupervised features successfully encode the essential information about the visual structures of the frames, the differences between the before and after frames can be minimally encoded. MID is implemented with a difference encoding module and a frame reconstruction module. The former tries to minimize the difference between frames, while the latter first maps the difference encoding to optical flow and then realizes frame reconstruction based on the learned optical flow. All the modules are jointly learned in an end-to-end way through the reconstruction loss. We verify the feasibility of MID in the image datasets by replacing the natural transformation of videos to artificially parameterized transformation for images. Experimental results show that MID is able to predict optical flow accurately based on the minimum difference coding and learn a powerful visual representation that greatly enhances other visual tasks.