Overview of the NLPCC 2024 Shared Task 7: Multi-lingual Medical Instructional Video Question Answering
摘要
Following the success of the 1-st Chinese Medical Instructional Video Question Answering (CMIVQA) Challenge in 2023, this year we have hosted the Multilingual Medical Instructional Video Question Answering (MMIVQA) shared task at the NLPCC 2024 conference. The MMIVQA task aims to promote the development of intelligent systems capable of understanding medical instructional video content and accurately delivering visual answers given the questions in a multilingual environment. This task encompasses three challenging tracks: (1) Multilingual Temporal Answer Grounding for Single Video (mTAGSV), (2) Multilingual Video Corpus Retrieval (mVCR), and (3) Multilingual Temporal Answer Grounding for Video Corpus (mTAGVC). These tracks cover different application scenarios ranging from single videos to large-scale video corpora, requiring participants’ systems to perform video content understanding, question answering, and temporal localization in both Chinese and English environments. We hope that this new MMIVQA challenge will provide more insights for first aid, medical emergencies, or medical education in multilingual settings.