Comparative performance of chatgpt and gemini in diagnostic classification and clinical reasoning for open-angle glaucoma: a standardized scenario-based study
摘要
To evaluate differences in performance between two large language models (LLMs), GPT-5.3 and Gemini 2.5 Pro, in diagnostic classification and clinical reasoning for primary open-angle glaucoma (POAG).
MethodsForty-eight guideline-based standardized cases were constructed. Using a unified prompt, we fed the cases into both models and obtained outputs on diagnosis, classification, and reasoning. A consensus of three glaucoma specialists served as the reference standard. We compared the diagnostic accuracy and classification consistency (Cohen’s κ) of the two models. Clinical reasoning ability was scored using a Likert scale based on logical coherence, evidence utilization, and conclusion consistency. We further analyzed error patterns and potential safety issues.
ResultsOverall diagnostic accuracy was 85.4% for GPT-5.3 and 75.0% for Gemini (P = 0.306). Both models performed well on typical cases but showed reduced accuracy on borderline cases, including ocular hypertension, suspected glaucoma, and early-stage POAG. For classification consistency, κ values were 0.675 for GPT-5.3 and 0.628 for Gemini. GPT-5.3 scored higher than Gemini in overall clinical reasoning (4.4 ± 0.6 vs. 3.9 ± 0.7, P = 0.011). Error pattern analysis indicated that Gemini was more prone to overdiagnosis and reasoning inconsistency, whereas GPT-5.3 was relatively conservative. Both models had low rates of unsafe outputs, though Gemini showed a slightly higher proportion.
ConclusionChatGPT and Gemini both demonstrate certain capabilities in diagnosing POAG, but their stability on borderline cases remains limited. Comparatively, GPT-5.3 shows higher consistency and more stable reasoning patterns. The application of LLMs in ophthalmic diagnostic support still requires cautious evaluation.