Following the success of the 1-st Chinese Medical Instructional Video Question Answering (CMIVQA) Challenge in 2023, this year we have hosted the Multilingual Medical Instructional Video Question Answering (MMIVQA) shared task at the NLPCC 2024 conference. The MMIVQA task aims to promote the development of intelligent systems capable of understanding medical instructional video content and accurately delivering visual answers given the questions in a multilingual environment. This task encompasses three challenging tracks: (1) Multilingual Temporal Answer Grounding for Single Video (mTAGSV), (2) Multilingual Video Corpus Retrieval (mVCR), and (3) Multilingual Temporal Answer Grounding for Video Corpus (mTAGVC). These tracks cover different application scenarios ranging from single videos to large-scale video corpora, requiring participants’ systems to perform video content understanding, question answering, and temporal localization in both Chinese and English environments. We hope that this new MMIVQA challenge will provide more insights for first aid, medical emergencies, or medical education in multilingual settings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Overview of the NLPCC 2024 Shared Task 7: Multi-lingual Medical Instructional Video Question Answering

  • Bin Li,
  • Yixuan Weng,
  • Qiya Song,
  • Lianhui Liang,
  • Xianwen Min,
  • Shoujun Zhou

摘要

Following the success of the 1-st Chinese Medical Instructional Video Question Answering (CMIVQA) Challenge in 2023, this year we have hosted the Multilingual Medical Instructional Video Question Answering (MMIVQA) shared task at the NLPCC 2024 conference. The MMIVQA task aims to promote the development of intelligent systems capable of understanding medical instructional video content and accurately delivering visual answers given the questions in a multilingual environment. This task encompasses three challenging tracks: (1) Multilingual Temporal Answer Grounding for Single Video (mTAGSV), (2) Multilingual Video Corpus Retrieval (mVCR), and (3) Multilingual Temporal Answer Grounding for Video Corpus (mTAGVC). These tracks cover different application scenarios ranging from single videos to large-scale video corpora, requiring participants’ systems to perform video content understanding, question answering, and temporal localization in both Chinese and English environments. We hope that this new MMIVQA challenge will provide more insights for first aid, medical emergencies, or medical education in multilingual settings.