错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CMGN: Cross-Modal Grounding Network for Temporal Sentence Retrieval in Video

  • Qun Zhang,
  • Bin Jiang,
  • Bolin Zhang,
  • Chao Yang

摘要

Temporal sentence grounding in video (TSGV) focuses on identifying the most pertinent temporal segment within an untrimmed video, given a natural language query. Its principal aim is to ascertain and retrieve the specific moment that impeccably aligns with the given query. Although the existing methods have done much research in this field and achieved specific achievements, there are still problems of massive calculation and insufficient grounding. Our method mainly focuses on obtaining better video and query features performing cross-modal feature fusion better, and locating more accurately when dealing with this problem. We propose an efficient Cross-Modal Grounding Network (CMGN) to balance the amount of computation and localization accuracy. In our proposed structure, we obtain the local context information through a bidirectional Gated Recurrent Unit (GRU). We obtain the start and end boundary characteristics for a better video presentation. Then, the two-channel structure, divided into a start channel and an end channel, captures the temporal relationships among several video segments sharing common boundaries. To validate the effectiveness of our method, extensive experiments were conducted on two datasets.