This paper presents the experiment scheme of Team IIEleven in the NLPCC 2024 shared task, Multilingual Medical Instructional Video Question Answering (MMIVQA) challenge, Track 3: Multilingual Temporal Answer Grounding in Video Corpus (mTAGVC). The objective of the mTAGVC task is to identify video spans most relevant to the given questions from a large multilingual medical instructional video corpus. We propose an Multilingual Visual-Textual Span Enhancement (MVTSE) method, which simultaneously performs two subtasks of video corpus retrieval and temporal answering grounding in video, and further strengthens the cross language ability and comparative learning performance of the model by subtitle text supplementation, multilingual data augmentation and hard sample selection method. The experimental results on the given dataset show that the proposed method achieves state-of-the-art (SOTA) performance, and ranks first in the mTAGVC track of the MMIVQA challenge.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multilingual Temporal Answer Grounding in Video Corpus with Enhanced Visual-Textual Integration

  • Tianxing Ma,
  • Yueyue Hu,
  • Shuang Jiang,
  • Zhenhao Yin,
  • Tianning Zang

摘要

This paper presents the experiment scheme of Team IIEleven in the NLPCC 2024 shared task, Multilingual Medical Instructional Video Question Answering (MMIVQA) challenge, Track 3: Multilingual Temporal Answer Grounding in Video Corpus (mTAGVC). The objective of the mTAGVC task is to identify video spans most relevant to the given questions from a large multilingual medical instructional video corpus. We propose an Multilingual Visual-Textual Span Enhancement (MVTSE) method, which simultaneously performs two subtasks of video corpus retrieval and temporal answering grounding in video, and further strengthens the cross language ability and comparative learning performance of the model by subtitle text supplementation, multilingual data augmentation and hard sample selection method. The experimental results on the given dataset show that the proposed method achieves state-of-the-art (SOTA) performance, and ranks first in the mTAGVC track of the MMIVQA challenge.