<p>With the explosion in the number of web videos, it has become a common practice to detect hot topics with web videos. However, each video clip contains multiple patterns, in which object actions might only appear in specific spatial areas or specific time periods, posing a huge challenge for web video hot topic detection. Fortunately, visual information during a specific time period and area will significantly enhance the rapid capture of key information, which is particularly important for detecting hot topics. Therefore, we propose a cross-modal associated learning method with spatial–temporal attention. It can automatically select discriminative time segments to detect hot topics by focusing on spatial regions with rich information. Firstly, after focusing on important keyframes related to the topic through temporal attention, spatial attention emphasizes the salient regions in the frame, thus incorporating discriminative features at the spatial level. Secondly, after integrating text structure knowledge into text semantic features, it can adaptively learn the weights of text features and visual features. Thirdly, adaptive learning of cross-modal fusion weights, achieving mutual guidance between text and visual information in attention dispersed association models to enhance feature learning. Finally, under the constraint of contrast loss, hot topics are detected with the similarity between features. Extensive experiments conducted on web videos from YouTube indicate that our method outperforms 8 leading state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-modal associated learning with spatial–temporal attention for hot topic detection

  • Chengde Zhang,
  • Shiyu Liu,
  • Xinyu Li,
  • Xia Xiao

摘要

With the explosion in the number of web videos, it has become a common practice to detect hot topics with web videos. However, each video clip contains multiple patterns, in which object actions might only appear in specific spatial areas or specific time periods, posing a huge challenge for web video hot topic detection. Fortunately, visual information during a specific time period and area will significantly enhance the rapid capture of key information, which is particularly important for detecting hot topics. Therefore, we propose a cross-modal associated learning method with spatial–temporal attention. It can automatically select discriminative time segments to detect hot topics by focusing on spatial regions with rich information. Firstly, after focusing on important keyframes related to the topic through temporal attention, spatial attention emphasizes the salient regions in the frame, thus incorporating discriminative features at the spatial level. Secondly, after integrating text structure knowledge into text semantic features, it can adaptively learn the weights of text features and visual features. Thirdly, adaptive learning of cross-modal fusion weights, achieving mutual guidance between text and visual information in attention dispersed association models to enhance feature learning. Finally, under the constraint of contrast loss, hot topics are detected with the similarity between features. Extensive experiments conducted on web videos from YouTube indicate that our method outperforms 8 leading state-of-the-art methods.