错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Overview of CHIP2023 Shared Task 4: CHIP-YIER Medical Large Language Model Evaluation

  • Han Hu,
  • Jun Yan,
  • Xiaozhen Zhang,
  • Zengtao Jiao,
  • Buzhou Tang

摘要

In recent years, large language models (LLMs) have demonstrated outstanding performance in natural language processing (NLP) tasks, showcasing remarkable capabilities in semantic understanding, text generation, and complex task reasoning. This emerging technology has also been applied to medical artificial intelligence, spanning traditional medical NLP tasks such as medical information extraction and medical entity normalization to practical applications in medical diagnosis, personalized treatment plan formulation, drug research and optimization, and early disease prediction. LLMs have achieved surprising results in these medical tasks. Major research institutions, universities, and companies have successively released their own large language models. However, due to the high risk characteristic of the medical industry, it is crucial to ensure their accuracy and reliability in clinical applications. To address this concern, the China Health Information Processing Conference (CHIP) introduced the evaluation task “CHIP-YIER Medical Large Model Evaluation” to assess the performance of large models in medical terminology, medical knowledge, adherence to clinical standards in diagnosis and treatment, and other medical aspects. The evaluation dataset is presented in the form of multiple-choice questions, comprising 1000 training set examples and 500 test set examples. Participating teams are required to select the correct answers based on the given descriptions. A total of 12 teams submitted valid results, with the highest F1 score of 0.7646.