<p>Large language models (LLMs) such as ChatGPT (GPT) and DeepSeek (DS) are increasingly explored in medical education and practice. This study compared the accuracy and concordance of GPT and DS in answering multiple-choice Progress Test questions (2013–2024) across three input formats: text-only, image-only, and full-question (vignette + image). GPT achieved 72.9% overall accuracy versus 65.1% for DS, with moderate inter-model agreement (κ = 0.591). Critically, for both models, full-question accuracy did not exceed text-only accuracy, indicating that the addition of visual content provided no global incremental benefit. Conversely, accuracy dropped significantly in the image-only condition, confirming that isolated visual information was less informative than textual history. Category-specific analyses revealed relative strengths for GPT in skin and body pictures and for DS in microscopic and radiographic images, though these did not translate into an overall multimodal advantage. Questions answered correctly had lower difficulty indices, while discrimination indices did not differ between correct and incorrect responses, and performance was similar across cognitive levels. These findings demonstrate that, under the tested conditions, LLM performance was predominantly text-driven rather than visuospatially integrated. Accordingly, ChatGPT and DeepSeek function as proficient text-based clinical reasoners, but not as effective multimodal diagnostic systems in this educational assessment context.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Differential Accuracy and Concordance of ChatGPT and DeepSeek for Image-Based Questions in Undergraduate Medical Education

  • Lorraine Silva Requena,
  • Joyce Santana Rizzi,
  • Zilda Maria Tosta Ribeiro,
  • Pedro Tadao Hamamoto Filho,
  • Renato Ferretti

摘要

Large language models (LLMs) such as ChatGPT (GPT) and DeepSeek (DS) are increasingly explored in medical education and practice. This study compared the accuracy and concordance of GPT and DS in answering multiple-choice Progress Test questions (2013–2024) across three input formats: text-only, image-only, and full-question (vignette + image). GPT achieved 72.9% overall accuracy versus 65.1% for DS, with moderate inter-model agreement (κ = 0.591). Critically, for both models, full-question accuracy did not exceed text-only accuracy, indicating that the addition of visual content provided no global incremental benefit. Conversely, accuracy dropped significantly in the image-only condition, confirming that isolated visual information was less informative than textual history. Category-specific analyses revealed relative strengths for GPT in skin and body pictures and for DS in microscopic and radiographic images, though these did not translate into an overall multimodal advantage. Questions answered correctly had lower difficulty indices, while discrimination indices did not differ between correct and incorrect responses, and performance was similar across cognitive levels. These findings demonstrate that, under the tested conditions, LLM performance was predominantly text-driven rather than visuospatially integrated. Accordingly, ChatGPT and DeepSeek function as proficient text-based clinical reasoners, but not as effective multimodal diagnostic systems in this educational assessment context.