<p>Purpose: To improve the explainability of a YOLOv5-based model for anterior segment disease diagnosis by combining gradient-weighted class activation mapping ++ (Grad-CAM++) and cut-and-paste validation and evaluate the influence of extracorneal information on diagnostic accuracy.</p><p>Methods: In total, 1039 slit-lamp photographs across nine diagnostic categories were analyzed using a previously developed YOLOv5 model. Grad-CAM++ was implemented to visualize the attention patterns across key network layers. To quantitatively assess clinical relevance, attention maps were compared against expert-delineated lesion boundaries using Intersection over Union (IoU). Furthermore, cut-and-paste validation was performed by systematically swapping the corneal regions with different backgrounds to probe the reliance of the model on contextual information.</p><p>Results: Grad-CAM++ analysis revealed hierarchical attention focus, progressing from broad in the early layers to highly localized in the final layers, with layer 23 showing distinct disease-specific patterns. Cut-and-paste validation demonstrated that model accuracy was highly dependent on the background context for certain diseases; for instance, the accuracy for infectious keratitis and acute primary angle closure dropped significantly upon background alteration. Critically, the clinical relevance of the attention maps measured by the IoU against expert annotations was significantly higher for correct predictions than for incorrect predictions, linking the model’s visual explanation to its diagnostic reliability.</p><p>Conclusion: Combining Grad-CAM++ with cut-and-paste validation provided a robust framework for evaluating the explainability of the YOLOv5 model. This dual approach reveals layer-specific attention dynamics and quantifies the model’s reliance on clinically relevant extracorneal features, thereby enhancing the transparency and trustworthiness of artificial intelligence-based diagnostic systems in ophthalmology.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

YOLOv5 Attention Analysis for Anterior Eye Disease Classification: Grad-CAM++ Feature Importance and Cut-and-Paste Validation

  • Yoshiyuki Kitaguchi,
  • Yuta Ueno,
  • Takefumi Yamaguchi,
  • Hiroki Maehara,
  • Dai Miyazaki,
  • Ryohei Nejima,
  • Takenori Inomata,
  • Naoko Kato,
  • Tai-ichiro Chikama,
  • Jun Ominato,
  • Tatsuya Yunoki,
  • Kinya Tsubota,
  • Masahiro Oda,
  • Kensaku Mori,
  • Yu Yoshinaga,
  • Rikako Iwasaki,
  • Kohji Nishida,
  • Tetsuro Oshika

摘要

Purpose: To improve the explainability of a YOLOv5-based model for anterior segment disease diagnosis by combining gradient-weighted class activation mapping ++ (Grad-CAM++) and cut-and-paste validation and evaluate the influence of extracorneal information on diagnostic accuracy.

Methods: In total, 1039 slit-lamp photographs across nine diagnostic categories were analyzed using a previously developed YOLOv5 model. Grad-CAM++ was implemented to visualize the attention patterns across key network layers. To quantitatively assess clinical relevance, attention maps were compared against expert-delineated lesion boundaries using Intersection over Union (IoU). Furthermore, cut-and-paste validation was performed by systematically swapping the corneal regions with different backgrounds to probe the reliance of the model on contextual information.

Results: Grad-CAM++ analysis revealed hierarchical attention focus, progressing from broad in the early layers to highly localized in the final layers, with layer 23 showing distinct disease-specific patterns. Cut-and-paste validation demonstrated that model accuracy was highly dependent on the background context for certain diseases; for instance, the accuracy for infectious keratitis and acute primary angle closure dropped significantly upon background alteration. Critically, the clinical relevance of the attention maps measured by the IoU against expert annotations was significantly higher for correct predictions than for incorrect predictions, linking the model’s visual explanation to its diagnostic reliability.

Conclusion: Combining Grad-CAM++ with cut-and-paste validation provided a robust framework for evaluating the explainability of the YOLOv5 model. This dual approach reveals layer-specific attention dynamics and quantifies the model’s reliance on clinically relevant extracorneal features, thereby enhancing the transparency and trustworthiness of artificial intelligence-based diagnostic systems in ophthalmology.