In the digital age, the widespread preference for instructional videos as an educational aid demands cutting-edge technologies that can swiftly pinpoint video segments in response to queries. In light of the likelihood that learners from various cultural backgrounds may pose questions in different languages, the multilingual temporal answer grounding in single video (mTAGSV) challenge has been introduced. MTAGSV requires models to precisely identify the particular time segment within the video that align with the query posed in Chinese or English. By addressing the limitations of existing monolingual approaches and their inefficacy with silent videos, we utilize optical character recognition (OCR) to enhance video information and leverage large language models (LLMs) to bridge linguistic gap. For videos that contain audio, the subtitles are extracted by an automated speech recognition (ASR) tool. For silent videos or those with insufficient audio-based subtitles, we leverage an OCR tool to extract textual content from video frames, refining the OCR-generated text to act as a surrogate for subtitles. Furthermore, we leverage LLMs to translate English queries into their Chinese equivalents, bridging the linguistic divide between queries and video contents. The MutualSL model is employed as the backbone network for extracting features from textual subtitles and visual frames. Through extensive experiments, we demonstrate that our proposed techniques enhance the task performance, securing first place in track 1 of NLPCC 2024 shared task 7.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Multilingual Temporal Answering Grounding in Single Video via LLM-Based Translation and OCR Enhancement

  • Huan Zhang,
  • Chen Zheng,
  • Yuanjing He,
  • Yan Zhao,
  • Yuxuan Lai

摘要

In the digital age, the widespread preference for instructional videos as an educational aid demands cutting-edge technologies that can swiftly pinpoint video segments in response to queries. In light of the likelihood that learners from various cultural backgrounds may pose questions in different languages, the multilingual temporal answer grounding in single video (mTAGSV) challenge has been introduced. MTAGSV requires models to precisely identify the particular time segment within the video that align with the query posed in Chinese or English. By addressing the limitations of existing monolingual approaches and their inefficacy with silent videos, we utilize optical character recognition (OCR) to enhance video information and leverage large language models (LLMs) to bridge linguistic gap. For videos that contain audio, the subtitles are extracted by an automated speech recognition (ASR) tool. For silent videos or those with insufficient audio-based subtitles, we leverage an OCR tool to extract textual content from video frames, refining the OCR-generated text to act as a surrogate for subtitles. Furthermore, we leverage LLMs to translate English queries into their Chinese equivalents, bridging the linguistic divide between queries and video contents. The MutualSL model is employed as the backbone network for extracting features from textual subtitles and visual frames. Through extensive experiments, we demonstrate that our proposed techniques enhance the task performance, securing first place in track 1 of NLPCC 2024 shared task 7.