错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Parameterized multi-perspective graph learning network for temporal sentence grounding in videos

  • Guangli Wu,
  • Zhijun Yang,
  • Jing Zhang

摘要

Temporal sentence grounding in videos (TSGV) aims to retrieve video segments from untrimmed videos that semantically matched a given query. Although existing methods have made significant progress in fine-grained intra- and inter-modal representations, they failed to comprehensively consider the redundancy of the entire video relative to the target segment and the fact that interactions could obscure crucial intra-modal information, which leads to the degradation of model performance. In this paper, we proposed a novel Parameterized Multi-Perspective Graph Learning Network for Temporal Sentence Grounding in Videos. Specifically, to effectively handle redundant information in video graphs, the concept of a parameterized network is introduced to dynamically construct new video graphs. Parameterizing the graph structure, making it adaptable to various video scenes while suppressing unnecessary redundant information. Furthermore, we designed a dual-path attention gating module that delves into cross-modal relationships while fully considering intra-modal information. The mechanism simultaneously considers the association between video and query from both inter- and intra-modal perspectives. This method allowed the model to better balance the local and global semantic consistency, further enhancing its representation capability for multimodal data. Extensive experiments on the ActivityNet Captions and Tacos benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art methods.