A video as acquired by a moving camera contains visual data pertaining to different places (i.e. a particular room, a corridor, etc.) that have been traversed through. Place proposal generation refers to delineating the correct place-related boundaries in the incoming video and encoding the respective visual data. This is an important problem since the resulting representation can be used for video-based place analysis. To this end, we propose a novel two-stage unsupervised place proposal generation framework that works real-time on unlabeled videos as acquired by a moving camera. First, each distinct place is delineated based on the continuous iterative partitioning of the incoming frames - considering their informativeness, coherency and plenitude. Following, “canonical scenes” within each generated place proposal are identified based on the hierarchical clustering of the respective frames. This enables larger or cluttered places that have multiple scenes to have better representations. Experimental results on benchmark data as well as real-time data demonstrate superior video place analysis performance as compared to a baseline approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unsupervised Real-Time Two-Stage Place Proposal Generation from a Moving Camera Video

  • H. Işıl Bozma

摘要

A video as acquired by a moving camera contains visual data pertaining to different places (i.e. a particular room, a corridor, etc.) that have been traversed through. Place proposal generation refers to delineating the correct place-related boundaries in the incoming video and encoding the respective visual data. This is an important problem since the resulting representation can be used for video-based place analysis. To this end, we propose a novel two-stage unsupervised place proposal generation framework that works real-time on unlabeled videos as acquired by a moving camera. First, each distinct place is delineated based on the continuous iterative partitioning of the incoming frames - considering their informativeness, coherency and plenitude. Following, “canonical scenes” within each generated place proposal are identified based on the hierarchical clustering of the respective frames. This enables larger or cluttered places that have multiple scenes to have better representations. Experimental results on benchmark data as well as real-time data demonstrate superior video place analysis performance as compared to a baseline approach.