<p>Retrieving video clips that match textual descriptions from untrimmed videos poses a significant challenge, often encountering issues of modality misalignment when integrating full-text information with clip or moment visual embeddings. To address this, we introduce the Multi-Hierarchical Semantic Graph Learning Network (MHSG), which aligns holistic visual data with comprehensive text information and segments visual details with keyword features, ensuring a more accurate cross-modal correspondence. MHSG leverages multi-hierarchical semantic structures to capture visual information at various levels of content granularity, aligning these extracted features with hierarchical text graphs based on inherent semantic relationships. Furthermore, we propose a transition to dense sampling and a local-global feature extraction strategy to improve clip prediction accuracy. Experiments demonstrate that MHSG achieves state-of-the-art performance on multiple benchmarks. Specifically, when using C3D features in the Charades-STA, ActivityNet Caption, and TACoS benchmarks, the model improves the R@1, IoU = 0.5 scores by 1.32%, 1.22%, and 0.12%, respectively. By achieving clear and explicit alignment through multi-hierarchical semantic structures, MHSG effectively enhances the retrieval of video clips matching textual descriptions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-hierarchical semantic graph learning for video moment retrieval

  • Xing Cheng,
  • De Han,
  • Nan Guo,
  • Xiaochun Ye,
  • Benjamin Rainer,
  • Peter Priller

摘要

Retrieving video clips that match textual descriptions from untrimmed videos poses a significant challenge, often encountering issues of modality misalignment when integrating full-text information with clip or moment visual embeddings. To address this, we introduce the Multi-Hierarchical Semantic Graph Learning Network (MHSG), which aligns holistic visual data with comprehensive text information and segments visual details with keyword features, ensuring a more accurate cross-modal correspondence. MHSG leverages multi-hierarchical semantic structures to capture visual information at various levels of content granularity, aligning these extracted features with hierarchical text graphs based on inherent semantic relationships. Furthermore, we propose a transition to dense sampling and a local-global feature extraction strategy to improve clip prediction accuracy. Experiments demonstrate that MHSG achieves state-of-the-art performance on multiple benchmarks. Specifically, when using C3D features in the Charades-STA, ActivityNet Caption, and TACoS benchmarks, the model improves the R@1, IoU = 0.5 scores by 1.32%, 1.22%, and 0.12%, respectively. By achieving clear and explicit alignment through multi-hierarchical semantic structures, MHSG effectively enhances the retrieval of video clips matching textual descriptions.