Widespread access to digital cameras has led to a significant increase of photo and video editing tools. Furthermore, sharing such content has led to the popularization of dozens of social media websites and apps. In this research we explore the extent to which a well-known foundational model, such as CLIP -Contrastive Language-Image Pre-Training- model, a pretrained model with semantic and scene recognition capabilities, can be used, without further training, as a moment retrieval searcher in video recordings. We propose two novel methods aimed at moment retrieval tasks in audiovisual data, namely VClipper-frame and VClipper-scene, and perform an empirical analysis on the effects of post-processing on the similarity vectors obtained from CLIP encoders. Our methods for zero-shot moment retrieval using CLIP scored better than the current state-of-the-art. Specifically, VClipper-frame reaches \(R@1_{0.5}=57.4\% \) and \(mAP_{0.5}=51.6\%\) . Compared to previous work, that reached values \(R@1_{0.5}=42.1\% \) and \(mAP_{0.5}=43.0\%\) , our approach represents an improvement of 15.3 points and 8.6 points, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VClipper: Moment Retrieval in Video Streams Using Zero-Shot and Context-Aware Foundational Models

  • Oriol Caravaca-Müller,
  • Joan Llobera,
  • Carles Ventura,
  • Ismael Benito-Altamirano

摘要

Widespread access to digital cameras has led to a significant increase of photo and video editing tools. Furthermore, sharing such content has led to the popularization of dozens of social media websites and apps. In this research we explore the extent to which a well-known foundational model, such as CLIP -Contrastive Language-Image Pre-Training- model, a pretrained model with semantic and scene recognition capabilities, can be used, without further training, as a moment retrieval searcher in video recordings. We propose two novel methods aimed at moment retrieval tasks in audiovisual data, namely VClipper-frame and VClipper-scene, and perform an empirical analysis on the effects of post-processing on the similarity vectors obtained from CLIP encoders. Our methods for zero-shot moment retrieval using CLIP scored better than the current state-of-the-art. Specifically, VClipper-frame reaches \(R@1_{0.5}=57.4\% \) and \(mAP_{0.5}=51.6\%\) . Compared to previous work, that reached values \(R@1_{0.5}=42.1\% \) and \(mAP_{0.5}=43.0\%\) , our approach represents an improvement of 15.3 points and 8.6 points, respectively.