Background <p>Ophthalmic diseases significantly impact vision and quality of life. Early diagnosis using fundus images is critical for timely treatment. Traditional deep learning models often lack accuracy, interpretability, and efficiency for multi-label classification tasks in ophthalmology.</p> Methods <p>We propose HAM-DNet, a hybrid deep learning model combining EfficientNetV2 and Vision Transformers (ViT) for multi-label ophthalmic disease detection. The model includes SE (Squeeze-and-Excitation) blocks for attention-based feature refinement and a U-Net-based lesion localization module for improved interpretability. The model was trained and tested on multiple fundus image datasets (ODIR-5&#xa0;K, Messidor, G1020, and Joint Shantou International Eye Centre).</p> Results <p>HAM-DNet achieved superior performance with an accuracy of 95.3%, precision of 96.2%, recall of 97.1%, AUC of 98.42, and F1-score of 96.75, while maintaining low computational cost (9.7 GFLOPS). It outperformed existing models including Shallow CNN and EfficientNet, particularly in handling multi-label classifications and reducing false positives and negatives.</p> Conclusions <p>HAM-DNet offers a robust, accurate, and interpretable solution for automated detection of multiple ophthalmic diseases. Its lightweight architecture makes it suitable for clinical deployment, especially in telemedicine and resource-constrained environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hybrid attention-based deep learning for multi-label ophthalmic disease detection on fundus images

  • Rabiya Hanfi,
  • Harsh Mathur,
  • Ritu Shrivastava

摘要

Background

Ophthalmic diseases significantly impact vision and quality of life. Early diagnosis using fundus images is critical for timely treatment. Traditional deep learning models often lack accuracy, interpretability, and efficiency for multi-label classification tasks in ophthalmology.

Methods

We propose HAM-DNet, a hybrid deep learning model combining EfficientNetV2 and Vision Transformers (ViT) for multi-label ophthalmic disease detection. The model includes SE (Squeeze-and-Excitation) blocks for attention-based feature refinement and a U-Net-based lesion localization module for improved interpretability. The model was trained and tested on multiple fundus image datasets (ODIR-5 K, Messidor, G1020, and Joint Shantou International Eye Centre).

Results

HAM-DNet achieved superior performance with an accuracy of 95.3%, precision of 96.2%, recall of 97.1%, AUC of 98.42, and F1-score of 96.75, while maintaining low computational cost (9.7 GFLOPS). It outperformed existing models including Shallow CNN and EfficientNet, particularly in handling multi-label classifications and reducing false positives and negatives.

Conclusions

HAM-DNet offers a robust, accurate, and interpretable solution for automated detection of multiple ophthalmic diseases. Its lightweight architecture makes it suitable for clinical deployment, especially in telemedicine and resource-constrained environments.