Purpose of Review <p>Artificial intelligence (AI) systems for retinal image interpretation have achieved diagnostic accuracy rivaling or exceeding that of human specialists across multiple conditions. The dominant deployment paradigm assumes that human clinicians provide meaningful oversight of AI output. This narrative review examines whether that assumption is valid by evaluating the published evidence on the quality and independence of human oversight in AI-assisted retinal diagnosis.</p> Recent Findings <p>Diabetic retinopathy (DR) screening programs increasingly rely on autonomous or semi-autonomous AI systems operating in high-throughput primary care settings, with non-specialist readers providing the human oversight layer. Published pivotal trial data show that FDA-cleared AI systems achieve 87–97% sensitivity for detecting more than mild DR, compared with 21% for general ophthalmologists and 47% for retina specialists by dilated ophthalmoscopy. These comparisons are clinically informative but not direct head-to-head measures of intrinsic diagnostic ability, as AI performance was assessed on fundus photographs against reading-center reference standards whereas the clinician estimates derive from dilated ophthalmoscopy. This performance gap, combined with evidence from the automation bias literature demonstrating that non-specialist reviewers are particularly susceptible to deference to algorithmic output, creates conditions favoring progressive erosion of independent human oversight. Critically, we did not identify any published retinal AI study measuring the correlation structure between human and AI errors, the parameter that determines whether human oversight adds independent safety value to the diagnostic ensemble.</p> Summary <p>Current evaluation frameworks assess retinal image AI accuracy and human-plus-AI accuracy but fail to capture whether human reviewers maintain independent diagnostic judgment. Without measuring error correlation and related oversight metrics, the field cannot distinguish between oversight that accurately catches AI failures and that which merely ratifies AI output. Notably, the complementary error profiles of AI systems (high sensitivity, lower specificity) and human clinicians (lower sensitivity, very high specificity) suggest substantial potential ensemble value, but only if human judgment remains independent. Longitudinal studies designed with error independence as a primary outcome are needed to fill this gap in retinal imaging.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Who Watches the Algorithm? The Unsolved Problem of Human Oversight Quality in Retinal Artificial Intelligence: A Narrative Review

  • Avnish A. Deobhakta,
  • Richard B. Rosen,
  • Nisha Chadha,
  • James C. Tsai,
  • Alon Harris,
  • Louis R. Pasquale

摘要

Purpose of Review

Artificial intelligence (AI) systems for retinal image interpretation have achieved diagnostic accuracy rivaling or exceeding that of human specialists across multiple conditions. The dominant deployment paradigm assumes that human clinicians provide meaningful oversight of AI output. This narrative review examines whether that assumption is valid by evaluating the published evidence on the quality and independence of human oversight in AI-assisted retinal diagnosis.

Recent Findings

Diabetic retinopathy (DR) screening programs increasingly rely on autonomous or semi-autonomous AI systems operating in high-throughput primary care settings, with non-specialist readers providing the human oversight layer. Published pivotal trial data show that FDA-cleared AI systems achieve 87–97% sensitivity for detecting more than mild DR, compared with 21% for general ophthalmologists and 47% for retina specialists by dilated ophthalmoscopy. These comparisons are clinically informative but not direct head-to-head measures of intrinsic diagnostic ability, as AI performance was assessed on fundus photographs against reading-center reference standards whereas the clinician estimates derive from dilated ophthalmoscopy. This performance gap, combined with evidence from the automation bias literature demonstrating that non-specialist reviewers are particularly susceptible to deference to algorithmic output, creates conditions favoring progressive erosion of independent human oversight. Critically, we did not identify any published retinal AI study measuring the correlation structure between human and AI errors, the parameter that determines whether human oversight adds independent safety value to the diagnostic ensemble.

Summary

Current evaluation frameworks assess retinal image AI accuracy and human-plus-AI accuracy but fail to capture whether human reviewers maintain independent diagnostic judgment. Without measuring error correlation and related oversight metrics, the field cannot distinguish between oversight that accurately catches AI failures and that which merely ratifies AI output. Notably, the complementary error profiles of AI systems (high sensitivity, lower specificity) and human clinicians (lower sensitivity, very high specificity) suggest substantial potential ensemble value, but only if human judgment remains independent. Longitudinal studies designed with error independence as a primary outcome are needed to fill this gap in retinal imaging.