Video moment localization, which aims to localize specific moments within a video based on a descriptive query, has gained significant attention in recent years. Despite substantial progress, most existing methods are supervised and rely on moment-level temporal annotations. In contrast, weakly-supervised approaches, which require only video-level annotations, remain relatively underexplored. In this chapter, we propose a novel end-to-end Siamese alignment network for weakly-supervised video moment retrieval. Specifically, we introduce a multi-scale Siamese module that progressively reduces the semantic gap between the visual and textual modalities. Furthermore, we propose a context-aware multiple instance learning module, which incorporates the influence of adjacent contexts, enhancing both moment-query and video-query alignment simultaneously. By optimizing both moment-level and video-level matching, our model effectively improves retrieval performance even with only weak video-level annotations. Extensive experiments on two benchmark datasets, ActivityNet Captions and Charades-STA, demonstrate the superior performance of our model over several state-of-the-art baselines.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Weakly-Supervised Video Moment Localization

  • Meng Liu,
  • Yupeng Hu,
  • Weili Guan,
  • Liqiang Nie

摘要

Video moment localization, which aims to localize specific moments within a video based on a descriptive query, has gained significant attention in recent years. Despite substantial progress, most existing methods are supervised and rely on moment-level temporal annotations. In contrast, weakly-supervised approaches, which require only video-level annotations, remain relatively underexplored. In this chapter, we propose a novel end-to-end Siamese alignment network for weakly-supervised video moment retrieval. Specifically, we introduce a multi-scale Siamese module that progressively reduces the semantic gap between the visual and textual modalities. Furthermore, we propose a context-aware multiple instance learning module, which incorporates the influence of adjacent contexts, enhancing both moment-query and video-query alignment simultaneously. By optimizing both moment-level and video-level matching, our model effectively improves retrieval performance even with only weak video-level annotations. Extensive experiments on two benchmark datasets, ActivityNet Captions and Charades-STA, demonstrate the superior performance of our model over several state-of-the-art baselines.