错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ST-CLIP: Spatio-Temporal Enhanced CLIP Towards Dense Video Captioning

  • Huimin Chen,
  • Pengfei Duan,
  • Mingru Huang,
  • Jingyi Guo,
  • Shengwu Xiong

摘要

Considering that Contrastive Language-Image Pretraining (CLIP) has demonstrated powerful capability of visual representation, many recent works directly apply CLIP to video understanding tasks. However, simply extending CLIP from the image domain to the video domain, treating video frames as static images, cannot effectively model the temporal information between video frames. In this paper, we propose a new factorized spatio-temporal self-attention paradigm to address inaccurate event descriptions caused by insufficient temporal relationship modeling between video frames in dense video captioning tasks. In addition, most existing dense video captioning methods mainly use visual content of videos and textual annotations, while ignoring the speech information which may provide more fine-grained complementarity beyond annotations. In this paper, we propose to transcribe speech into text as auxiliary data to enhance visual representation. Extensive experiments on the YouCook2 and ViTT demonstrate that our method achieves state-of-the-art performance.