Parameterized multi-perspective graph learning network for temporal sentence grounding in videos
摘要
Temporal sentence grounding in videos (TSGV) aims to retrieve video segments from untrimmed videos that semantically matched a given query. Although existing methods have made significant progress in fine-grained intra- and inter-modal representations, they failed to comprehensively consider the redundancy of the entire video relative to the target segment and the fact that interactions could obscure crucial intra-modal information, which leads to the degradation of model performance. In this paper, we proposed a novel Parameterized Multi-Perspective Graph Learning Network for Temporal Sentence Grounding in Videos. Specifically, to effectively handle redundant information in video graphs, the concept of a parameterized network is introduced to dynamically construct new video graphs. Parameterizing the graph structure, making it adaptable to various video scenes while suppressing unnecessary redundant information. Furthermore, we designed a dual-path attention gating module that delves into cross-modal relationships while fully considering intra-modal information. The mechanism simultaneously considers the association between video and query from both inter- and intra-modal perspectives. This method allowed the model to better balance the local and global semantic consistency, further enhancing its representation capability for multimodal data. Extensive experiments on the ActivityNet Captions and Tacos benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art methods.