With the rapid proliferation of surveillance and web videos, video moment localization has become a critical task in video content analysis, attracting significant attention from both academia and industry. However, this task presents several challenges, including effective temporal context modeling, intelligent generation of moment candidates, and ensuring efficiency and scalability for real-world applications. To address these challenges, we propose a novel deep end-to-end cross-modal hashing network. Specifically, we introduce a video encoder based on a bidirectional temporal convolutional network, which simultaneously generates moment candidates and learns their representations. By capturing temporal contextual structures across multiple scales, the video encoder produces enriched moment representations. Complementing this, we design an independent query encoder to comprehensively understand user intent. We then present a cross-modal hashing module that projects the heterogeneous video and query representations into a shared isomorphic Hamming space, enabling compact hash code learning. The relevance of each “moment-query” pair is estimated efficiently via the Hamming distance. Our model not only ensures effectiveness in localization but also achieves notable efficiency and scalability, as the hash codes can be learned offline. Extensive experiments on public datasets demonstrate the superiority of our method compared to several state-of-the-art approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient Hashing Based Video Moment Localization

  • Meng Liu,
  • Yupeng Hu,
  • Weili Guan,
  • Liqiang Nie

摘要

With the rapid proliferation of surveillance and web videos, video moment localization has become a critical task in video content analysis, attracting significant attention from both academia and industry. However, this task presents several challenges, including effective temporal context modeling, intelligent generation of moment candidates, and ensuring efficiency and scalability for real-world applications. To address these challenges, we propose a novel deep end-to-end cross-modal hashing network. Specifically, we introduce a video encoder based on a bidirectional temporal convolutional network, which simultaneously generates moment candidates and learns their representations. By capturing temporal contextual structures across multiple scales, the video encoder produces enriched moment representations. Complementing this, we design an independent query encoder to comprehensively understand user intent. We then present a cross-modal hashing module that projects the heterogeneous video and query representations into a shared isomorphic Hamming space, enabling compact hash code learning. The relevance of each “moment-query” pair is estimated efficiently via the Hamming distance. Our model not only ensures effectiveness in localization but also achieves notable efficiency and scalability, as the hash codes can be learned offline. Extensive experiments on public datasets demonstrate the superiority of our method compared to several state-of-the-art approaches.