CMGN: Cross-Modal Grounding Network for Temporal Sentence Retrieval in Video
摘要
Temporal sentence grounding in video (TSGV) focuses on identifying the most pertinent temporal segment within an untrimmed video, given a natural language query. Its principal aim is to ascertain and retrieve the specific moment that impeccably aligns with the given query. Although the existing methods have done much research in this field and achieved specific achievements, there are still problems of massive calculation and insufficient grounding. Our method mainly focuses on obtaining better video and query features performing cross-modal feature fusion better, and locating more accurately when dealing with this problem. We propose an efficient Cross-Modal Grounding Network (CMGN) to balance the amount of computation and localization accuracy. In our proposed structure, we obtain the local context information through a bidirectional Gated Recurrent Unit (GRU). We obtain the start and end boundary characteristics for a better video presentation. Then, the two-channel structure, divided into a start channel and an end channel, captures the temporal relationships among several video segments sharing common boundaries. To validate the effectiveness of our method, extensive experiments were conducted on two datasets.