Importance <p>Artificial intelligence (AI), machine learning (ML), and deep learning (DL) have already produced clinically deployed ophthalmic systems for diabetic retinopathy (DR) screening and high-performing algorithms for image-level classification and lesion segmentation [1-8]. However, clinical retinal interpretation is not equivalent to assigning a global disease label. A clinically useful reader must associate lesions with anatomical locations, compare the same location across time, and express the finding in a way that can guide treatment or referral.</p> Objective <p>To review evidence that the major unsolved problem for present ophthalmic AI is not whether an image can be globally classified as DR, but whether vision-language models (VLMs), large vision-language models (LVLMs), and multimodal large language models (MLLMs) can perform grounded lesion-location-change reasoning at the level required for clinical care [9-25].</p> Evidence Review <p>We performed a critical narrative review of peer-reviewed ophthalmology, medical imaging, and AI literature, supplemented by regulatory documents. The review focused on DR, diabetic macular edema (DME), hard exudates, fundus foundation models, medical VLMs, visual grounding, hallucination, and reinforcement learning [6-7, 14-39]. Non-peer-reviewed news and trade-magazine sources were removed from the evidentiary base.</p> Findings <p>Image-level and screening systems have reached high performance in selected settings, as shown by JAMA and FDA-cleared autonomous DR work [4-7]. Pixel-level lesion detection can also be achieved when lesion masks and anatomical labels are available, as shown by IDRiD and DeepDR [8, 40-41]. In contrast, current VLMs perform inconsistently on ophthalmic multimodal image analysis and remain vulnerable to spatial-reasoning errors, visual hallucination, and weak anatomical localization [14-25]. The clinical example of'increased hard exudates in the macula' requires foveal/macular localization, hard-exudate detection, registration across visits, quantification, and text generation as a single grounded chain [40-45]. A reinforcement learning (RL) approach cannot solve this by rewarding final disease labels; it requires verifiable dense rewards, including fovea-localization error, lesion Dice/IoU, lesion-to-fovea distance, and longitudinal change metrics [24-25, 35-36].</p> Conclusions and Relevance <p>The next clinically decisive frontier for ophthalmic AI is grounded clinical reading: the ability of a VLM to bind lesion identity, anatomical location, and longitudinal change into a verifiable statement. This review proposes lesion-location-change reasoning as the critical benchmark for future ophthalmic VLMs.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From Diabetic Retinopathy Classification to Grounded Clinical Reading: Lesion-location-change Reasoning as the Unresolved Challenge for Ophthalmic Vision-language Models

  • Hitoshi Tabuchi,
  • Mao Tanabe,
  • Keita Kihara,
  • Takaaki Moriguchi,
  • Fumi Gomi

摘要

Importance

Artificial intelligence (AI), machine learning (ML), and deep learning (DL) have already produced clinically deployed ophthalmic systems for diabetic retinopathy (DR) screening and high-performing algorithms for image-level classification and lesion segmentation [1-8]. However, clinical retinal interpretation is not equivalent to assigning a global disease label. A clinically useful reader must associate lesions with anatomical locations, compare the same location across time, and express the finding in a way that can guide treatment or referral.

Objective

To review evidence that the major unsolved problem for present ophthalmic AI is not whether an image can be globally classified as DR, but whether vision-language models (VLMs), large vision-language models (LVLMs), and multimodal large language models (MLLMs) can perform grounded lesion-location-change reasoning at the level required for clinical care [9-25].

Evidence Review

We performed a critical narrative review of peer-reviewed ophthalmology, medical imaging, and AI literature, supplemented by regulatory documents. The review focused on DR, diabetic macular edema (DME), hard exudates, fundus foundation models, medical VLMs, visual grounding, hallucination, and reinforcement learning [6-7, 14-39]. Non-peer-reviewed news and trade-magazine sources were removed from the evidentiary base.

Findings

Image-level and screening systems have reached high performance in selected settings, as shown by JAMA and FDA-cleared autonomous DR work [4-7]. Pixel-level lesion detection can also be achieved when lesion masks and anatomical labels are available, as shown by IDRiD and DeepDR [8, 40-41]. In contrast, current VLMs perform inconsistently on ophthalmic multimodal image analysis and remain vulnerable to spatial-reasoning errors, visual hallucination, and weak anatomical localization [14-25]. The clinical example of'increased hard exudates in the macula' requires foveal/macular localization, hard-exudate detection, registration across visits, quantification, and text generation as a single grounded chain [40-45]. A reinforcement learning (RL) approach cannot solve this by rewarding final disease labels; it requires verifiable dense rewards, including fovea-localization error, lesion Dice/IoU, lesion-to-fovea distance, and longitudinal change metrics [24-25, 35-36].

Conclusions and Relevance

The next clinically decisive frontier for ophthalmic AI is grounded clinical reading: the ability of a VLM to bind lesion identity, anatomical location, and longitudinal change into a verifiable statement. This review proposes lesion-location-change reasoning as the critical benchmark for future ophthalmic VLMs.