<p>In this study, we present <b>MedS-Bench</b>, a comprehensive benchmark to evaluate large language models (LLMs) in clinical contexts, <b>MedS-Bench</b>, spanning 11 high-level clinical tasks. We evaluate nine leading LLMs, <i>e.g</i>., MEDITRON, Llama 3, Mistral, GPT-4, Claude-3.5, <i>etc</i>. and found that most models struggle with these complex tasks. To address these limitations, we developed <b>MedS-Ins</b>, a large-scale instruction-tuning dataset for medicine. <b>MedS-Ins</b> comprises 58 medically oriented language corpora, totaling 5M instances with 19K instructions, across 122 tasks. To demonstrate the dataset’s utility, we conducted a proof-of-concept experiment by performing instruction tuning on a lightweight, open-source medical language model. The resulting model, <b>MMedIns-Llama 3</b>, significantly outperformed existing models on various clinical tasks. To promote further advancements, we have made <b>MedS-Ins</b> fully accessible and invite the research community to contribute to its expansion. Additionally, we have launched a dynamic leaderboard for <b>MedS-Bench</b>, to track the development progress of medical LLMs.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards evaluating and building versatile large language models for medicine

  • Chaoyi Wu,
  • Pengcheng Qiu,
  • Jinxin Liu,
  • Hongfei Gu,
  • Na Li,
  • Ya Zhang,
  • Yanfeng Wang,
  • Weidi Xie

摘要

In this study, we present MedS-Bench, a comprehensive benchmark to evaluate large language models (LLMs) in clinical contexts, MedS-Bench, spanning 11 high-level clinical tasks. We evaluate nine leading LLMs, e.g., MEDITRON, Llama 3, Mistral, GPT-4, Claude-3.5, etc. and found that most models struggle with these complex tasks. To address these limitations, we developed MedS-Ins, a large-scale instruction-tuning dataset for medicine. MedS-Ins comprises 58 medically oriented language corpora, totaling 5M instances with 19K instructions, across 122 tasks. To demonstrate the dataset’s utility, we conducted a proof-of-concept experiment by performing instruction tuning on a lightweight, open-source medical language model. The resulting model, MMedIns-Llama 3, significantly outperformed existing models on various clinical tasks. To promote further advancements, we have made MedS-Ins fully accessible and invite the research community to contribute to its expansion. Additionally, we have launched a dynamic leaderboard for MedS-Bench, to track the development progress of medical LLMs.