Background <p>Large language models (LLMs), such as ChatGPT, have demonstrated promising potential in medical knowledge retrieval and clinical decision support. DeepSeek, a China-developed model released in 2025, has been proposed as a medical AI tool, but its performance in healthcare settings remains underexplored.</p> Methods <p>We systematically evaluated DeepSeek’s performance in urolithiasis through two approaches. First, we measured its accuracy and consistency using 157 single-choice questions from public datasets and the Chinese National Medical Licensing Examination. Second, we compared the clinical decision-making of DeepSeek and ChatGPT using three real-world urolithiasis cases. Responses were evaluated against those of clinicians across four dimensions: readability, medical knowledge accuracy, diagnostic appropriateness, and logical coherence.</p> Results <p>In the medical knowledge task, DeepSeek achieved an accuracy above 83%, comparable to ChatGPT, with no significant difference in readability. However, in simulated clinical scenarios, DeepSeek underperformed in diagnostic reasoning and in avoiding unnecessary testing. The DeepSeek-R1 (R1) model scored significantly lower than both ChatGPT-o3 (R3) and physicians across several dimensions.</p> Conclusion <p>DeepSeek shows strong potential in structured medical knowledge retrieval but remains limited in its ability to support clinical decision-making. With continued model refinement, it may serve as a valuable tool in medical education and clinical practice.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Systematic evaluation of deepseek in urolithiasis: from medical knowledge to clinical decision support

  • Anguo Zhao,
  • Rongkang Li,
  • Lei Peng,
  • Rui Liang,
  • Ruonan Sun,
  • Fan Wu,
  • Zhengyan Wang,
  • Xiaojian Xu,
  • Jun Zhang,
  • Jianquan Hou

摘要

Background

Large language models (LLMs), such as ChatGPT, have demonstrated promising potential in medical knowledge retrieval and clinical decision support. DeepSeek, a China-developed model released in 2025, has been proposed as a medical AI tool, but its performance in healthcare settings remains underexplored.

Methods

We systematically evaluated DeepSeek’s performance in urolithiasis through two approaches. First, we measured its accuracy and consistency using 157 single-choice questions from public datasets and the Chinese National Medical Licensing Examination. Second, we compared the clinical decision-making of DeepSeek and ChatGPT using three real-world urolithiasis cases. Responses were evaluated against those of clinicians across four dimensions: readability, medical knowledge accuracy, diagnostic appropriateness, and logical coherence.

Results

In the medical knowledge task, DeepSeek achieved an accuracy above 83%, comparable to ChatGPT, with no significant difference in readability. However, in simulated clinical scenarios, DeepSeek underperformed in diagnostic reasoning and in avoiding unnecessary testing. The DeepSeek-R1 (R1) model scored significantly lower than both ChatGPT-o3 (R3) and physicians across several dimensions.

Conclusion

DeepSeek shows strong potential in structured medical knowledge retrieval but remains limited in its ability to support clinical decision-making. With continued model refinement, it may serve as a valuable tool in medical education and clinical practice.