A multi-task benchmark dataset for assessing large language models in traditional Chinese medicine
摘要
Traditional Chinese Medicine (TCM) is a holistic medical system with millennia of accumulated clinical experience, playing a vital role in global healthcare. However, its implicit reasoning, diverse textual forms, and lack of standardization pose major challenges for computational modeling and assessment. Large Language Models (LLMs) have demonstrated remarkable potential in processing natural language across general medical domains, yet their systematic assessment in the TCM domain remains underdeveloped. Existing benchmarks either focus narrowly on factual question answering or lack domain-specific tasks and clinical realism. To fill this gap, we introduce Multi-Task TCM Assessment Benchmark (MTCMB), a multi-task benchmark dataset for assessing LLMs on TCM knowledge, reasoning, and safety. Developed in collaboration with certified TCM experts, MTCMB comprises 12 sub-datasets spanning five major categories: knowledge question answering (QA), language understanding, diagnostic reasoning, prescription recommendation, and safety assessment. The benchmark integrates real-world case records, national licensing exams, and authoritative textbooks to provide an authentic testbed. To demonstrate the utility of the resource, we provide a technical validation using representative LLMs, confirming the dataset’s ability to discriminate between model capabilities across varying difficulty levels. We release MTCMB as an open-access resource to foster the development of competent and trustworthy medical AI systems.