<p>Traditional Chinese Medicine (TCM) is a holistic medical system with millennia of accumulated clinical experience, playing a vital role in global healthcare. However, its implicit reasoning, diverse textual forms, and lack of standardization pose major challenges for computational modeling and assessment. Large Language Models (LLMs) have demonstrated remarkable potential in processing natural language across general medical domains, yet their systematic assessment in the TCM domain remains underdeveloped. Existing benchmarks either focus narrowly on factual question answering or lack domain-specific tasks and clinical realism. To fill this gap, we introduce Multi-Task TCM Assessment Benchmark (MTCMB), a multi-task benchmark dataset for assessing LLMs on TCM knowledge, reasoning, and safety. Developed in collaboration with certified TCM experts, MTCMB comprises 12 sub-datasets spanning five major categories: knowledge question answering (QA), language understanding, diagnostic reasoning, prescription recommendation, and safety assessment. The benchmark integrates real-world case records, national licensing exams, and authoritative textbooks to provide an authentic testbed. To demonstrate the utility of the resource, we provide a technical validation using representative LLMs, confirming the dataset’s ability to discriminate between model capabilities across varying difficulty levels. We release MTCMB as an open-access resource to foster the development of competent and trustworthy medical AI systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A multi-task benchmark dataset for assessing large language models in traditional Chinese medicine

  • Shufeng Kong,
  • Yuanyuan Wei,
  • Xingru Yang,
  • Zijie Wang,
  • Hao Tang,
  • Jiuqi Qin,
  • Shuting Lan,
  • Nuan Cui,
  • Lei Gao,
  • Yingheng Wang,
  • Junwen Bai,
  • Zhuangbin Chen,
  • Zibin Zheng,
  • Caihua Liu,
  • Hao Liang

摘要

Traditional Chinese Medicine (TCM) is a holistic medical system with millennia of accumulated clinical experience, playing a vital role in global healthcare. However, its implicit reasoning, diverse textual forms, and lack of standardization pose major challenges for computational modeling and assessment. Large Language Models (LLMs) have demonstrated remarkable potential in processing natural language across general medical domains, yet their systematic assessment in the TCM domain remains underdeveloped. Existing benchmarks either focus narrowly on factual question answering or lack domain-specific tasks and clinical realism. To fill this gap, we introduce Multi-Task TCM Assessment Benchmark (MTCMB), a multi-task benchmark dataset for assessing LLMs on TCM knowledge, reasoning, and safety. Developed in collaboration with certified TCM experts, MTCMB comprises 12 sub-datasets spanning five major categories: knowledge question answering (QA), language understanding, diagnostic reasoning, prescription recommendation, and safety assessment. The benchmark integrates real-world case records, national licensing exams, and authoritative textbooks to provide an authentic testbed. To demonstrate the utility of the resource, we provide a technical validation using representative LLMs, confirming the dataset’s ability to discriminate between model capabilities across varying difficulty levels. We release MTCMB as an open-access resource to foster the development of competent and trustworthy medical AI systems.