<p>The recent study by Wu et al. (2025) comparing DeepSeek-R1 and ChatGPT-4o on the Chinese National Medical Licensing Examination (CNMLE) provides an important contribution to understanding large language model (LLM) performance in non-English medical contexts. While their findings highlight the potential of LLMs in medical knowledge assessment, several methodological issues merit further discussion. First, the exclusive use of Chinese-language items without bilingual comparison may favor DeepSeek-R1, which demonstrates strong performance in Chinese, over ChatGPT-4o, whose training corpus is predominantly English-based. Second, the evaluation was conducted before the release of GPT-5, leading to potential disparities in reasoning capabilities between models. Third, the restriction to multiple-choice questions limits the assessment to factual recall rather than higher-order reasoning or clinical judgment. We commend the authors for initiating this valuable cross-linguistic analysis and suggest that future studies incorporate bilingual testing, ensure model functional parity, and include open-ended clinical items to more comprehensively evaluate LLMs’ reasoning and interpretive competence in real-world medical education contexts.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards A Fair Duel: Reflections on the Evaluation of DeepSeek-R1 and ChatGPT-4o in Chinese Medical Education

  • Shangxuan Li

摘要

The recent study by Wu et al. (2025) comparing DeepSeek-R1 and ChatGPT-4o on the Chinese National Medical Licensing Examination (CNMLE) provides an important contribution to understanding large language model (LLM) performance in non-English medical contexts. While their findings highlight the potential of LLMs in medical knowledge assessment, several methodological issues merit further discussion. First, the exclusive use of Chinese-language items without bilingual comparison may favor DeepSeek-R1, which demonstrates strong performance in Chinese, over ChatGPT-4o, whose training corpus is predominantly English-based. Second, the evaluation was conducted before the release of GPT-5, leading to potential disparities in reasoning capabilities between models. Third, the restriction to multiple-choice questions limits the assessment to factual recall rather than higher-order reasoning or clinical judgment. We commend the authors for initiating this valuable cross-linguistic analysis and suggest that future studies incorporate bilingual testing, ensure model functional parity, and include open-ended clinical items to more comprehensively evaluate LLMs’ reasoning and interpretive competence in real-world medical education contexts.