Background <p>To evaluate the clinical applicability of three generic Vision-Large-Language Models (VLLMs) — OpenAI’s GPT-4omni, GPT-4V(ision) and Google’s Gemini in detecting and diagnosing inherited retinal diseases (IRDs), using fundus photographs.</p> Methods <p>The head-to-head comparative study curated 60 ultra-widefield (UWF) fundus images of 30 IRD patients from the National University Hospital, Singapore. Additionally, ten normal, open-sourced UWF fundus images were included for comparison. The 70 fundus images were analysed by the three VLLMs using standardised prompts to generate descriptions of 10 specified retinal features and provide clinical insights. Each VLLM received 2100 scores for descriptions across ten features, rated by three blinded consultant-level graders using three-point scale (0 = poor, 1 = borderline, 2 = good). Clinical insights including disease detection, diagnosis and pathological gene inference evaluated against clinical ground-truth.</p> Results <p>GPT-4o achieved the highest mean quality score in feature description (1.64 [0.697], mean [SEM]), outperforming GPT-4V (1.57 [0.738]) and Gemini (1.46 [0.800]; both <i>p</i> &lt; 0.001). All models demonstrated high detection accuracy (<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41433_2025_4013_Article_IEq1.gif" Format="GIF" Height="15" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\ge\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>≥</mo> </math></EquationSource> </InlineEquation>81.4%), but Gemini incorrectly classified all normal fundus images as IRD. GPT-4omni (65.7%) outperformed GPT-4V (50%) and Gemini (60%) in diagnosis accuracy. Gene inference precision remained low (<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41433_2025_4013_Article_IEq2.gif" Format="GIF" Height="15" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\le\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>≤</mo> </math></EquationSource> </InlineEquation>20.3%) across all models. High concordance was observed across all models between feature descriptions and diagnoses (<InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41433_2025_4013_Article_IEq1.gif" Format="GIF" Height="15" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\ge\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>≥</mo> </math></EquationSource> </InlineEquation>97.1%), between diagnoses and clinical recommendations (100%).</p> Conclusions <p>GPT-4omni and GPT-4V demonstrated promising potential in detecting IRDs from fundus photographs, with good feature extraction capabilities and high detection accuracy. Gemini struggled with misidentifying normal fundus images. All three VLLMs require further refinement to improve diagnostic accuracy and gene inference.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative analysis of generic vision-language models in detecting and diagnosing inherited retinal diseases using fundus photographs

  • Xiang Meng,
  • Wendy Meihua Wong,
  • Krithi Pushpanathan,
  • Sahana Srinivasan,
  • Cancan Xue,
  • Meng Wang,
  • Heng Miao,
  • Hengtong Li,
  • Liping Yang,
  • Ling-Ping Cen,
  • Li Jia Chen,
  • Hwei Wuen Chan,
  • Yih-Chung Tham,
  • Ching-Yu Cheng

摘要

Background

To evaluate the clinical applicability of three generic Vision-Large-Language Models (VLLMs) — OpenAI’s GPT-4omni, GPT-4V(ision) and Google’s Gemini in detecting and diagnosing inherited retinal diseases (IRDs), using fundus photographs.

Methods

The head-to-head comparative study curated 60 ultra-widefield (UWF) fundus images of 30 IRD patients from the National University Hospital, Singapore. Additionally, ten normal, open-sourced UWF fundus images were included for comparison. The 70 fundus images were analysed by the three VLLMs using standardised prompts to generate descriptions of 10 specified retinal features and provide clinical insights. Each VLLM received 2100 scores for descriptions across ten features, rated by three blinded consultant-level graders using three-point scale (0 = poor, 1 = borderline, 2 = good). Clinical insights including disease detection, diagnosis and pathological gene inference evaluated against clinical ground-truth.

Results

GPT-4o achieved the highest mean quality score in feature description (1.64 [0.697], mean [SEM]), outperforming GPT-4V (1.57 [0.738]) and Gemini (1.46 [0.800]; both p < 0.001). All models demonstrated high detection accuracy ( \(\ge\) 81.4%), but Gemini incorrectly classified all normal fundus images as IRD. GPT-4omni (65.7%) outperformed GPT-4V (50%) and Gemini (60%) in diagnosis accuracy. Gene inference precision remained low ( \(\le\) 20.3%) across all models. High concordance was observed across all models between feature descriptions and diagnoses ( \(\ge\) 97.1%), between diagnoses and clinical recommendations (100%).

Conclusions

GPT-4omni and GPT-4V demonstrated promising potential in detecting IRDs from fundus photographs, with good feature extraction capabilities and high detection accuracy. Gemini struggled with misidentifying normal fundus images. All three VLLMs require further refinement to improve diagnostic accuracy and gene inference.