TCM-Eval: A Multi-dimensional Benchmark Framework for Evaluating Large Language Models in Traditional Chinese Medicine
摘要
Recently, domain-specific large language models (LLMs) have been rapidly developed and prosperous in many specialized fields, yet their performance in traditional Chinese medicine (TCM) remains unclear because of the lack of a reliable and user-friendly benchmarking framework. In this paper, we propose TCM-Eval, an automated, scalable, and open-source system for evaluating LLMs on TCM scenarios. In TCM-Eval, we design six core dimensional tasks modules plus 1 integrated module, and standardize a total of 6,813 questions covering 29 specialized data sources. Furthermore, we integrate 16 LLMs, including the popular commercial, open-source, and state-of-the-art proprietary LLMs, into TCM-Eval. To provide a fine-grained and systematic assessment, we also design 4 multi-dimensional metrics across 7 types of questions. Our framework shapes a new standard for TCM LLM evaluation, offering a robust foundation for the relevant studies and clinical applications.