Clinically aligned artificial intelligence for glaucoma diagnosis: enhancing retinal nerve fibre layer interpretation from fundus images
摘要
This study aimed to develop an interpretable artificial intelligence (AI) screening system that replicates a specialist’s evaluation of fundus photographs. The system analyses three signs: visible retinal nerve fibre layer (RNFL) defects (as observed on fundus photography, not OCT-based measurement), vertical cup-to-disc ratio (VCDR) and rim-to-disc ratio (RDR).
Subjects/methodsA total of 773 fundus images from the independent test cohort were annotated by three fellowship-trained glaucoma specialists, followed by repeated consensus meetings to improve annotation consistency. Two models were trained on the development cohort: an EfficientNet-B4 classifier for RNFL and a U-Net with an EfficientNet-B4 encoder for optic disc/cup (OD/OC) segmentation. During the testing phase, a referral was triggered whenever a feature was classified as abnormal, with the system explicitly reporting the specific sign that prompted the referral decision.
ResultsAmong the 773 expert-annotated cases, 749 were included in the analysis after excluding low-quality or insufficient-information images (AI drop rate: 3.1%). The final independent test cohort comprised 268 referral and 481 non-referral cases, providing a basis for evaluation. The combined model achieved a sensitivity of 0.903 (95% confidence interval [CI], 0.868–0.938), specificity of 0.821 (95% CI, 0.787–0.855) and a positive predictive value (PPV) of 0.738 (95% CI, 0.690–0.785). The system captured complementary aspects of glaucomatous optic neuropathy that are often missed by single-feature approaches.
ConclusionsConsensus-based annotation, combined with lesion-level modelling, enhances alignment with clinical reasoning. By providing an explicit referral rationale, the system fosters trust in AI-assisted glaucoma screening and facilitates adoption in clinical settings.