Accurate diagnosis of ocular diseases is crucial for early detection and effective treatment. While recent vision-language interaction methods have achieved significant performance in medical diagnosis, existing approaches rely solely on convolutional neural networks for visual feature extraction, overlooking critical diagnostic textual data, and individual-specific traits. To address these limitations, this paper proposes an innovative framework that integrates visual, diagnostic semantic, and generative knowledge for robust ocular disease classification. The framework consists of three key modules: a Scalable Diagnostic Information Network (SDIN), a Heterogeneous Graph Attention Network (HGAT), and a Fusion Decision Module (FDM). SDIN combines binocular fundus images, diagnostic keywords, and external information generated by a Large Language Model into a comprehensive graph-based representation. HGAT employs a dual-level attention mechanism to enhance feature representation across heterogeneous modalities, while FDM effectively integrates features from binocular images for accurate multi-label ocular disease classification. Additionally, patient-specific information, such as age and gender, is encoded to further refine the feature representation. Experiments conducted on the OIA-ODIR dataset demonstrate that the proposed model achieves state-of-the-art performance, with a 1.19% improvement in final score over previous methods, showcasing its effectiveness and robustness in diverse clinical scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ocular Disease Classification Based on Heterogeneous Interaction Among Visual, Diagnostic Semantics, and Generative Knowledge

  • Zechang Xiong,
  • Zhenyan Ji,
  • Jiuqian Dai,
  • Hui Liu,
  • Wenhui Chen,
  • Shen Yin,
  • Jose Enrique Armendariz-Inigo

摘要

Accurate diagnosis of ocular diseases is crucial for early detection and effective treatment. While recent vision-language interaction methods have achieved significant performance in medical diagnosis, existing approaches rely solely on convolutional neural networks for visual feature extraction, overlooking critical diagnostic textual data, and individual-specific traits. To address these limitations, this paper proposes an innovative framework that integrates visual, diagnostic semantic, and generative knowledge for robust ocular disease classification. The framework consists of three key modules: a Scalable Diagnostic Information Network (SDIN), a Heterogeneous Graph Attention Network (HGAT), and a Fusion Decision Module (FDM). SDIN combines binocular fundus images, diagnostic keywords, and external information generated by a Large Language Model into a comprehensive graph-based representation. HGAT employs a dual-level attention mechanism to enhance feature representation across heterogeneous modalities, while FDM effectively integrates features from binocular images for accurate multi-label ocular disease classification. Additionally, patient-specific information, such as age and gender, is encoded to further refine the feature representation. Experiments conducted on the OIA-ODIR dataset demonstrate that the proposed model achieves state-of-the-art performance, with a 1.19% improvement in final score over previous methods, showcasing its effectiveness and robustness in diverse clinical scenarios.