Background <p>DeepSeek-R1, an open-source reasoning large language model (LLM) clinically deployed in Chinese hospitals, still lacks validation in ophthalmology.</p> Aims <p>To compare DeepSeek-R1 against OpenAI’s o1 and upgraded o3 models in diagnostic accuracy and reasoning capability across diverse ophthalmic conditions.</p> Methods <p>We evaluated 98 standardized case vignettes covering13 ophthalmic sub-specialties, each supplied with an expert-validated diagnostic hierarchy, differential list, and reasoning chain. Model performance was assessed with a diagnosis matrix focused on final-diagnosis (FDx) accuracy; incorrect outputs were resubmitted with key diagnostic clues (reasoning-augmented, RA prompt) to test self-correction. Reasoning capacity was quantified by the number/score of diagnostic clues retrieved per case across 13 predefined domains.</p> Results <p>DeepSeek-R1 achieved an 87.8% FDx accuracy, comparable to o3 (91.8%, P = .34) and higher than o1 (58.2%, <i>P</i> &lt; .001). Similar trends were observed for others accuracy (global <i>P</i> &lt; .001). Agreement was moderate–high between R1 and o3 (κ = 0.42–1.00), but slight with o1 (κ = 0.12–0.32). R1 and o3 identified more diagnostic clues than o1 (median count = 4 vs. 3, median score = 100 vs. 80; <i>P</i> &lt; .001). RA prompts corrected 50.0%, 62.5% and 41.5% of FDx errors for R1, o3, and o1, raising FDx accuracy to 93.9%, 96.9%, and 80.6% respectively.</p> Conclusions <p>DeepSeek-R1 matched o3 and outperformed o1 in diagnostic accuracy and reasoning, retrieving nearly all expert-defined clues. Its open-source nature, low cost and strong performance support its use as a practical aid for ophthalmic decision-making.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of DeepSeek-R1 for Ophthalmic Diagnosis and Reasoning: A Comparison with OpenAI o1 and o3

  • Shuai Ming,
  • Xi Yao,
  • Qingge Guo,
  • Dandan Chen,
  • Xiaohong Guo,
  • Kunpeng Xie,
  • Bo Lei

摘要

Background

DeepSeek-R1, an open-source reasoning large language model (LLM) clinically deployed in Chinese hospitals, still lacks validation in ophthalmology.

Aims

To compare DeepSeek-R1 against OpenAI’s o1 and upgraded o3 models in diagnostic accuracy and reasoning capability across diverse ophthalmic conditions.

Methods

We evaluated 98 standardized case vignettes covering13 ophthalmic sub-specialties, each supplied with an expert-validated diagnostic hierarchy, differential list, and reasoning chain. Model performance was assessed with a diagnosis matrix focused on final-diagnosis (FDx) accuracy; incorrect outputs were resubmitted with key diagnostic clues (reasoning-augmented, RA prompt) to test self-correction. Reasoning capacity was quantified by the number/score of diagnostic clues retrieved per case across 13 predefined domains.

Results

DeepSeek-R1 achieved an 87.8% FDx accuracy, comparable to o3 (91.8%, P = .34) and higher than o1 (58.2%, P < .001). Similar trends were observed for others accuracy (global P < .001). Agreement was moderate–high between R1 and o3 (κ = 0.42–1.00), but slight with o1 (κ = 0.12–0.32). R1 and o3 identified more diagnostic clues than o1 (median count = 4 vs. 3, median score = 100 vs. 80; P < .001). RA prompts corrected 50.0%, 62.5% and 41.5% of FDx errors for R1, o3, and o1, raising FDx accuracy to 93.9%, 96.9%, and 80.6% respectively.

Conclusions

DeepSeek-R1 matched o3 and outperformed o1 in diagnostic accuracy and reasoning, retrieving nearly all expert-defined clues. Its open-source nature, low cost and strong performance support its use as a practical aid for ophthalmic decision-making.