Multilingual Temporal Answer Grounding in Video Corpus with Enhanced Visual-Textual Integration
摘要
This paper presents the experiment scheme of Team IIEleven in the NLPCC 2024 shared task, Multilingual Medical Instructional Video Question Answering (MMIVQA) challenge, Track 3: Multilingual Temporal Answer Grounding in Video Corpus (mTAGVC). The objective of the mTAGVC task is to identify video spans most relevant to the given questions from a large multilingual medical instructional video corpus. We propose an Multilingual Visual-Textual Span Enhancement (MVTSE) method, which simultaneously performs two subtasks of video corpus retrieval and temporal answering grounding in video, and further strengthens the cross language ability and comparative learning performance of the model by subtitle text supplementation, multilingual data augmentation and hard sample selection method. The experimental results on the given dataset show that the proposed method achieves state-of-the-art (SOTA) performance, and ranks first in the mTAGVC track of the MMIVQA challenge.